Choosing the Right Luma Model Before You Waste a Generation

Luma puts a dozen AI models behind one board; this guide shows which one to pick and how to prompt it right.

By RamthaMedia

RamthaMedia Free eBooks  ·  August 2026

Price: Priceless
 ·  15 min read

Preface

Luma puts a dozen AI image and video models — Nano Banana, GPT Image 2, Seedream, Ray3.14, Ray3.2, Seedance 2.5, Veo, Sora, Kling — behind one board and one chat window, and that is exactly the problem: which one, for which shot, prompted which way? This book walks through picking the right model before you write a word, restyling footage you already shot, turning a flat image into editable pieces, and saving a workflow so you never rebuild it by hand again.

Chapter 1

One Login, Nine Models, One Deadline Tomorrow

A freelance producer has three deliverables due before a client call tomorrow morning: a hero product shot, a fifteen-second product video, and a set of layered social crops in four aspect ratios. Two years ago that meant three different subscriptions, three different accounts, and three different sets of habits to remember. Tonight it means opening one board.

Luma is built around that single board rather than around any one model. The toolbar's Photo icon opens image generation and hands you a choice between models like Nano Banana, Nano Banana Pro, GPT Image 2 and GPT Image 1.5. The Play icon opens video generation and offers Ray3.14, Ray3, Ray 3.2, Seedance 2.5, Veo, Sora and Kling. A Waveform icon opens speech, music and sound effects. Every asset you generate lands on the same board, next to everything else.

The agent chat sits alongside the toolbar and can be used instead of it: describing what you want in plain language, uploading a reference file, or selecting an existing asset and asking for a variation. Both paths reach the same models. The toolbar suits someone who already knows which tool they want; the chat suits someone who wants to describe the outcome and let the agent suggest the route there.

None of that solves the producer's actual problem, though. Nine models sitting behind one login is not the same as knowing which one to open first. A hero shot, a video and four social crops are three different jobs, and treating them as one prompt fired at whichever model happens to be selected is how a deadline gets missed one wasted generation at a time.

The chapters that follow work through that decision in the order a real job usually forces it: pick the image model for the job, pick the video model for the shot, write the prompt in the shape that model actually rewards, then move into the tools – restyling, layering, saving a workflow – that turn one good result into something repeatable.

What you can actually do here

Luma's toolbar opens the same three panels no matter which model sits behind them. The table below is a starting point for matching the job to the model before spending a generation on the wrong one.

Picking an image model

Use Who it fits Where Worth knowing
Fast, cheap exploration with readable in-image text Anyone drafting concepts before a client sees them Toolbar → Photo icon → select Nano Banana → prompt Fast iteration, strong text rendering
Capped near 1 megapixel per generation
Client-facing, high-resolution final artwork Agencies delivering finished campaign assets Toolbar → Photo icon → select Nano Banana Pro → set 4K Resolution control up to 4K, strong typography
More resource-intensive than the base model
Complex, multi-element scenes composed correctly on the first try Anyone building infographics, UI mockups or layered ads Toolbar → Photo icon → select GPT Image 2 → follow the 5-part prompt Plans structure before drawing; up to 4 references
No transparent background support
Transparent-background assets for compositing elsewhere Designers who need to drop a subject onto another layout Toolbar → Photo icon → select GPT Image 1.5 Transparency support GPT Image 2 lacks

Picking a video model

Use Who it fits Where Worth knowing
Default, fastest general-purpose video generation Anyone producing routine social or product clips Toolbar → Play icon → select Ray3.14 Native 1080p, HDR, EXR export, seamless looping
No character reference support
Keeping one specific character's face consistent across shots Anyone building a narrative or brand mascot series Toolbar → Play icon → select Ray3 → upload character reference Character reference works across T2V, I2V and V2V
Slower than Ray3.14
Restyling a video you already shot, at its original length Producers localising or re-skinning an existing edit Right-click a video asset → Modify Video → Ray 3.2 Up to 64 keyframes, output matches source duration exactly
Needs a prompt or keyframes – source video alone errors
Longer, multi-beat scenes with timestamped control Anyone directing a sequence rather than one shot Toolbar → Play icon → select Seedance 2.5 → write timecoded beats Stronger continuity across cuts, up to 20s duration
Needs stable references to hold multi-person identity

Production tools

Use Who it fits Where Worth knowing
Splitting a finished poster or ad into swappable pieces Anyone who needs to update one element without rebuilding the whole design Right-click image → Extract Layers → describe what to separate Each element keeps a transparent background and stays editable
Over-fragmenting creates more layers than you can manage
Turning a workflow you tuned once into something you can run again Anyone repeating a brand treatment, reskin or format across many assets Select the asset or prompt → ask the agent to turn it into a Skill Captures the prompt, model choice, defaults and preservation rules together
Built-in Skills must be forked before they can be edited

Chapter 2

Picking an Image Model Before You Write a Word

Nano Banana is the model to open first for almost anything exploratory. It generates quickly, renders readable text inside an image better than most alternatives, and holds a character or style steady across several generations – useful when a concept needs three or four quick variations before anyone commits to a direction. Its ceiling is real, though: output is capped near one megapixel across its ten preset aspect ratios, and it will often fail to render a recognisable likeness of a real person unless you supply a reference image of them first.

Nano Banana Pro is the same lineage scaled up for delivery rather than drafting. Resolution control runs from a fast 1K up to a full 4K, it accepts up to fourteen reference images against Nano Banana's own smaller limit, and its text rendering carries into brand typography work cleanly enough for client-facing output. The cost is that it is more resource-intensive, and – like its smaller sibling – it can still return a generic result for a niche or unusual art style.

GPT Image 2 solves a different problem: complex, multi-element scenes that need to be composed correctly rather than merely rendered attractively. It plans the structure of an image before generating it, which is why it handles infographics, layered advertisements and UI mockups noticeably better than models built around straightforward keyword matching. It accepts up to four reference images, renders text with better than 95% accuracy across scripts including Latin, CJK, Arabic, Hindi and Bengali, and reaches native 2K with an optional 4K upscale. What it will not do is produce a transparent background – for anything that needs to be composited onto another layout afterward, that rules it out entirely.

That gap is exactly what GPT Image 1.5 exists to cover. Where the newer model trades transparency for stronger reasoning about complex scenes, the older one keeps RGBA support. Reaching for GPT Image 1.5 specifically for a transparent product cutout, rather than assuming the newest model does everything the older one did, is the kind of small routing decision that saves a generation rather than wasting one.

The pattern worth carrying forward is that speed, resolution, text accuracy, reference count and transparency all trade against each other differently in each model. There is no single 'best' image model in Luma – there is a fastest one, a highest-resolution one, a most-structurally-reliable one and a most-compositable one, and the job in front of you decides which property you actually need this time.

Chapter 3

Picking a Video Model for the Shot You Actually Need

Ray3.14 is the default recommended video model, and it earns that position on speed and range rather than any single standout feature: native 1080p with HDR, an EXR export option for professional colour grading, six aspect ratios from portrait to ultrawide, and seamless looping for product showcases. It supports start-frame-and-end-frame keyframes for precise interpolation between two images. What it will not do is hold a character's face steady across separate generations, and it carries no native audio.

Ray3 exists specifically to cover that first gap. It carries the same keyframe and aspect-ratio support as Ray3.14, but adds character reference: upload an image of a person and the model works to keep that identity consistent across text-to-video, image-to-video and video-to-video generations alike. The trade is speed – Ray3 runs slower than Ray3.14, so it is worth reaching for only once character consistency is actually the requirement, not by default.

Ray 3.2 is a different kind of tool entirely: it does not generate a new video from a prompt, it restyles a video you already have. The source clip is required; the output always matches the source's duration exactly, up to a maximum of twenty seconds, and there is no separate duration setting to override that. Up to sixty-four keyframes can be anchored at exact frame indexes in the source timeline, which is what makes it useful for art-directed restyles rather than a single blanket filter over the whole clip.

Seedance 2.5 is built for scenes rather than isolated clips. It accepts timestamped beats inside one prompt – describing what happens at 0-5 seconds, 5-10 seconds and so on – and holds continuity across cuts more reliably than treating every shot as its own generation. It also handles multi-person scenes more practically when each character is given a stable, named role inside the prompt rather than left ambiguous.

The field guide behind these models includes a decision flowchart built around exactly this kind of question: does the job need consistency and temporal stability, character consistency across shots, native audio, high-energy multi-subject action, the longest possible single clip, or the ability to modify existing footage. Answering that question before opening the toolbar is what actually saves a generation – not memorising every model's spec sheet.

Chapter 4

Writing Prompts the Models Actually Listen To

Every model in Luma rewards specificity over adjectives, but the shape of that specificity differs by model, and writing one universal prompt style across all of them is the single most common way to waste a generation. Ray3.14 wants mid-action verbs and present tense – 'running', not 'begins to run' – plus the secondary consequences of motion: wind in hair, dust kicked up, fabric moving. It is a positive-only model; negative prompting inside the prompt itself tends to work against you rather than for you.

GPT Image 2 follows what its own documentation calls the Golden Order: background and environment first, then the subject, then specific details, then constraints – moving from wide context to narrow specifics, because that mirrors how its reasoning engine parses the request. Its editing mode wants the opposite of a polite description: terse, direct commands stating exactly what to change and what to keep, with no flowery justification attached.

Seedance 2.0 works from a six-ingredient formula: a specific subject, one concrete action, the environment, one camera movement, a lighting source and mood, and a style. The most common mistake is asking a four-second clip to carry the complexity that only a fifteen-second one can hold – one subject, one action and one camera movement is the right scope for a short clip; a longer one can support a structured, timecoded sequence of several beats.

Across every one of these models, one habit degrades output the same way: reaching for vague aesthetic words instead of concrete visual facts. 'Stunning', 'cinematic' and 'high quality' describe a feeling, not a scene, and a model's reasoning has nothing to attach them to. 'Overcast daylight, shallow depth of field' or 'anamorphic 2.39:1, teal-and-orange grade, lens flare' gives it something to actually render.

None of this is really about vocabulary. It is about giving each model the specific kind of information it was built to act on – a mid-action verb here, a wide-to-narrow structure there, a concrete visual fact instead of a mood word everywhere – and that difference is worth more to the final result than any amount of extra prompt length.

You may also like:
Getting the Most From Snapgen Before the Credits Run Out

Chapter 5

Restyling Footage You Already Shot

Ray 3.2 starts from the assumption that you already have something worth keeping: a camera move, a performance, a product angle, a timed edit. The job is not to invent a shot from nothing but to change how that shot looks while its motion, framing and timing stay exactly where they were. That single distinction changes how a prompt for it should be written.

The prompt should describe the target end state, never the transformation process. Writing 'change the sky to purple' invites the model to think about the act of changing something over time; writing 'a purple sky over the mountain landscape' simply tells it what the final frame should contain. Negation works against this model for the same reason it works against Ray3.14 – 'remove the people from the street' should become 'an empty cobblestone street at dawn', describing the result rather than the subtraction.

Keyframes are the deepest control surface here. A keyframe is a still image pinned to one exact frame index in the source video, and up to sixty-four of them can be supplied across a single clip. A common production pattern is exporting a handful of frames from the source footage, editing them in third-party software or with one of Luma's own image models, and feeding them back in at their original frame indexes – the model then interpolates the restyled look across everything in between.

An Enhance toggle sits on by default and can rewrite a rough instruction into stronger V2V-style guidance before the run happens. That is useful for a fast, half-formed idea. It is the wrong setting for a carefully worded production prompt that has already been tested and approved, because Enhance will still alter it before it reaches the model – turning it off is what keeps an approved prompt exactly as written.

Ray 3.2's Motion, Structure and Character controls decide how much of the source's movement, shapes and performance carry through independently of each other. The practical habit worth adopting is starting with the lightest setup, watching what the result is missing, and adding exactly one control to fix that specific gap – rather than pushing every slider up at once and losing the ability to tell which one caused the change.

Chapter 6

From Draft to Delivery: Resolution, HDR and Color Grading

Ray3.14 and Ray 3.2 both expose four resolution tiers, from a rough 360p draft up through 540p and 720p to a final 1080p. Working a job in the lowest usable tier during exploration and only moving to 1080p for the version that is actually going out the door is not a shortcut – it is the difference between spending a generation cheaply on a decision that might get thrown away and spending it expensively on one that has already been approved.

HDR is off by default in both models and exists for a specific reason: dramatic lighting scenarios – sunsets, neon signage, stage lighting, fire – carry a wider dynamic range than standard output captures. Turning it on is worth doing for premium brand work, automotive or fashion delivery, and modern OLED-targeted output; it is unnecessary weight for a quick social clip that nobody will view on a display capable of showing the difference.

EXR export sits behind HDR and is easy to miss because it lives as a secondary option rather than a headline feature. Once enabled, it attaches a frame sequence with full linear floating-point colour data to the output, in a form that finishing tools such as Nuke, Resolve, Flame or Baselight can actually use. A video destined for a professional colour-grading pipeline needs this; a video destined for a social feed does not, and turning it on for the latter only adds render time for a file nobody downstream will open.

A separate inference-mode choice – Quality or Speed – governs how the render itself is prioritised. Quality mode is the right default for anything final. Speed mode trades some of that for faster turnaround, and it is the correct choice specifically during exploration, quick review rounds, or batch testing many directions before choosing one to finish properly.

Put together, a reliable production pattern looks like this: explore in Speed mode at a low resolution tier without HDR, pick the direction that works, then re-render the chosen take in Quality mode at 1080p with HDR and EXR switched on only if a colourist downstream actually needs the file that way.

You may also like:
What Higgsfield Runs Under One Login

Chapter 7

Turning One Flat Image into Editable Pieces

A finished poster, product ad or social graphic is normally treated as one object: change anything, regenerate everything. Layers exists to break that constraint. Right-clicking any image on the board and choosing Extract Layers opens a window where you describe what should be separated, and each extracted element comes back with a transparent background as its own editable piece inside a layer stack.

The instruction you give can be broad or exact, and which one to use depends on how much you already know. A broad instruction – 'separate this design into the most useful layers for editing and rearranging' – suits an unfamiliar image where you are not yet sure what needs to change. An exact instruction – naming the headline, the product, its shadow, the logo and the background individually, and specifying which elements should stay grouped together – suits a production-ready job where you already know exactly what will be swapped.

Once extracted, each layer can be moved independently in the composition view, reordered in the stack to change what sits in front of what, grouped with related layers using a keyboard shortcut, and shown or hidden without being deleted. That last distinction matters more than it looks: hiding a layer keeps it available for comparison later, while deleting it removes that option. The safer habit during any uncertain edit is to hide first and delete only once a direction is confirmed.

Modifying an individual layer works the same way as any other image prompt, scoped to just that element – 'replace the red running shoe with a white leather sneaker photographed from the same angle, preserving the existing lighting and shadow' changes only the shoe, leaving the background, the logo and the copy exactly as they were. The regenerated version is added as a new layer alongside the original rather than replacing it outright, which means both states stay available at once.

The practical value of Layers is not that it makes prettier output than a fresh generation would. It is that it scopes the change to the one element that actually needs to move, while everything that was already working – the layout, the typography, the lighting – stays untouched.

Chapter 8

Saving a Workflow So You Never Rebuild It

A Skill is a saved creative recipe rather than a single output: the prompt, the model choice, the defaults, the references and the preservation rules that made a particular workflow work, captured once so it can be run again on a new asset without reconstructing all of it from memory.

There is no single correct way to create one. A prompt already written as a text asset on the board can be turned into a Skill directly. A plain chat instruction – 'create a Skill that repaints any object in a matte colour' – can be converted the same way. A finished output that carries its own generation footprint can be inspected and turned back into the recipe that produced it, and a whole board process, with its intermediate steps and reference examples, can be distilled into one Skill if the board is organised clearly enough to read.

Running a saved Skill is as simple as selecting the asset to work on, typing @ in chat to bring up the Skill by name, and adding whatever the Skill still needs as input. Custom instructions can sit on top of a run without replacing the Skill itself – 'apply the gold material, and make the eyes emeralds' still follows the Skill's locked workflow while adjusting one specific detail for that run alone. Presets built into a Skill let a common variation be triggered by name rather than rewritten each time.

Skills that are built-in or shared by someone else are read-only until forked. Forking creates an editable copy, and any later change – a different default aspect ratio, an added preset, a rewritten prompt body – happens on that copy without touching the original. A Skill can also be shared between boards through a link, which carries the whole workflow with it rather than a written description someone else has to reconstruct.

The signal worth watching for is repetition, not perfection. A workflow that has been useful exactly once is worth keeping as an output. A workflow that has already been useful twice is worth saving as a Skill, because the second and third times it is needed, the cost of rebuilding it from memory is what a Skill removes entirely.

Chapter 9

What Luma Won't Let You Do, and Where Your Work Goes

Luma's terms restrict several things a producer might otherwise assume were fair game. Reverse engineering, decompiling or attempting to access the non-public parts of the Services is prohibited outright. Publishing benchmarks or performance information about the Services is specifically forbidden, not merely discouraged. Using the Services or their output on a service-bureau or rental basis – reselling access to someone else rather than using it yourself – is restricted unless a specific exception in an order permits it.

Rate limits, request quotas and throughput caps can be imposed, modified and enforced at Luma's discretion, and attempting to work around them through multiple accounts or credential sharing is treated as a breach rather than a workaround. If a limit is exceeded, access can be throttled, suspended or terminated without prior notice, and restoring it depends on the cause being fixed to Luma's satisfaction – not on time simply passing.

API credentials are the user's own responsibility to protect; any compromise must be reported within twenty-four hours of becoming aware of it, and Luma carries no liability for unauthorised use of credentials that were not adequately safeguarded. Accounts inactive for an extended period can be terminated outright, which is worth knowing before treating an old account as a safe place to park unfinished work.

On the data side, conversations with the agent – including the images, video and text submitted as input – are collected, and that content may be reproduced inside the generated output itself. Payment information is handled by a third-party processor rather than stored directly. Users under thirteen are not permitted to use the Services at all, and anyone between thirteen and eighteen needs a parent or guardian's consent to be bound by the agreement.

None of this changes what the models can create. It changes what a producer is responsible for once they start creating with them: keeping credentials secure, staying inside whatever rate limit applies, and treating anything typed into the agent chat as content that becomes part of the record rather than a private draft.

Questions readers actually ask

Which Luma model should I use if I don't know yet which one I need?

Nano Banana for images and Ray3.14 for video are the two defaults documented as the starting point for most work – fast, versatile, and good enough to judge a direction before switching to a more specialised model.

Can Luma keep the same character's face consistent across several video generations?

Yes, but only through Ray3, which supports a character reference image across text-to-video, image-to-video and video-to-video generations. Ray3.14 does not support this at all.

Why did my Ray 3.2 restyle come out longer or shorter than I expected?

It can't – Ray 3.2's output always matches its source video's exact duration, up to a maximum of twenty seconds. There is no separate duration setting to change that.

Do I need a prompt to run Ray 3.2, or can I use keyframes alone?

Either works. A prompt alone creates a pure prompt-driven restyle, keyframes alone create a visual-reference-driven one, and both together give the strongest control. A source video with neither a prompt nor keyframes will error.

What's the difference between hiding and deleting a layer in the Layers editor?

Hiding removes a layer from the visible composition without discarding it, so it can be brought back for comparison. Deleting removes it permanently. The safer habit during an uncertain edit is to hide first.

Can I edit a Skill that was shared with me by someone else?

Not directly if it's marked read-only. Fork it first to create your own editable copy – the original Skill and anyone else using it are unaffected by changes made to the fork.

Why does GPT Image 2 refuse to generate a transparent background?

It doesn't support transparency at all – that capability sits with GPT Image 1.5 instead. For any asset that needs to be composited onto another layout, GPT Image 1.5 is the model to choose, not the newer one.

What happens if I exceed a rate limit on Luma's API?

Access can be throttled, suspended or terminated without prior notice, at Luma's discretion, and there's no guaranteed restoration timeline – it depends on the cause of the excess usage being addressed first.

Is there an age restriction on using Luma?

Users under 13 are not authorised to use the Services at all. Users between 13 and 18 need a parent or legal guardian's consent to be bound by the agreement.

Does turning on HDR automatically give me an EXR export?

No – HDR and EXR export are separate toggles, but EXR specifically requires HDR to already be enabled. Turning HDR off after enabling EXR can cause the EXR option to disappear from the panel.

Can I reuse a Skill I built on one board on a different board?

Yes – open the Skills library, choose the Skill, and use Share to generate a link. Importing that link on another board adds the whole Skill, including its saved defaults and presets.

Will content I type into Luma's agent chat stay private?

The conversation content, including any images, video or text submitted as input, is collected, and that content may be reproduced inside the generated output itself – it is not treated as a private draft.

Contact / More useful information from RamthaMedia

    Official source links:
    Luma

    As an Amazon Associate, RamthaMedia earns from qualifying purchases.


    Disclaimer: This eBook is compiled from publicly available information and was accurate at the time of writing. For full and up-to-date details, please visit the official website linked above. RamthaMedia accepts no legal liability for any decision made on the basis of this eBook, and nothing here is professional, financial or legal advice. The image used for the cover page is illustrative only – a stock photo from Pexels or an AI-generated image, never a real photograph of the site described.

    RamthaMedia
    RamthaMedia

    About the Founder – A. Ravinder
    A. Ravinder is the Founder, Author, Digital Publisher, and Editor-in-Chief of RamthaMedia, a Telugu-focused digital media and publishing platform dedicated to delivering trusted news, practical knowledge, books, and smart buying guides.
    With strong experience in digital publishing, journalism, content research, and affiliate product analysis, he creates reliable, easy-to-understand, and value-driven content that helps readers make informed decisions in their daily lives.
    Through RamthaMedia, he combines news reporting, book publishing, educational resources, and honest product reviews — building a trusted knowledge ecosystem for Telugu and Indian audiences.

    Articles: 293