By RamthaMedia
RamthaMedia Free eBooks · August 2026
Price: Priceless
· 16 min read
Preface
Leonardo puts dozens of image and video models, a character-locking system, and a production-grade upscaler behind one login — which is exactly why new arrivals waste hours guessing. This book walks through choosing the right model for a shot, keeping a character's face the same across forty generations, reading the token economy before you commit to a plan, and turning a rough idea into an image sharp enough for a billboard.
Chapter 1
One Login, an Absurd Number of Models
Priya has one afternoon to turn a client's product photo into a rotating showcase clip, a matching packaging mockup, and three social posts, and she has never opened Leonardo before. She logs in expecting one text box and one button. What she gets is a wall of names — Lucid Origin, Nano Banana Pro, Seedream 5.0 Pro, FLUX 2 Pro, Kling O3 Omni Video, Veo 3.1, Hailuo 2.3 — and no obvious reason to pick one over another.
The wall exists because no single AI model does everything well. A model tuned to render photorealistic skin and cinematic lighting is not the same model that renders razor-sharp text on a poster, and neither is built to hold a character's face steady across twenty seconds of motion. Most AI platforms solve this by picking one model and living with its limits. Leonardo solves it differently: it puts dozens of image and video engines, built by different labs, behind one login and one shared token balance, and lets the job decide which one gets used.
That is also why the model list keeps changing shape. Sora 2 was available on the platform through most of 2026 and was pulled in July after a provider-side change at OpenAI — nothing Leonardo did, just a licensing arrangement that ended. Veo 3.1 and Seedance 2.5 stepped into the gap the same week. A model list built this way is never a finished catalogue; it is closer to a rotating shelf, worth checking before starting a project that depends on one specific model's look.
For someone who does not want to learn the roster, Auto mode reads the prompt and picks a plausible match — useful for a first pass, less useful once the work has a specific look in mind. Reading the list itself is simpler once it is grouped by job rather than by name. Product photography and architectural renders lean toward models built for composition and lighting control. Marketing copy with legible text on it — posters, packaging, banners — leans toward models built specifically for text rendering. Motion work splits again: fast, cheap models for testing an idea; a heavier model with synced audio for the final cinematic pass.
None of this is obvious from the pricing page, which lists token costs but not which model suits which job — that pairing only shows up once the sample galleries and model descriptions are actually compared side by side. A freelancer working across several client briefs a week saves real time by keeping a private shortlist: one model for photoreal product shots, one for text-heavy graphics, one for quick video drafts, one for the polished final render.
Priya's afternoon works out because the product shot, the packaging mockup and the rotating clip are three different jobs wearing one login. Picking the model per job, rather than hoping one model handles all three, is the first real skill this platform asks for — and it is the one almost nobody arrives already knowing.
Chapter 2
What a Token Actually Buys You
Before Dev commits his agency's card to a plan, he wants to know one plain thing: how many usable images and videos does that money actually buy in a month. The pricing page does not answer this directly. It answers in tokens, and a token is not the same size twice — a single image might cost a few dozen, a single video clip can cost thousands, because the two jobs demand wildly different amounts of computing power underneath.
This is why Leonardo prices in a currency instead of a flat 'unlimited' promise. A flat price for every kind of output would either force the platform to ration video invisibly or charge everyone for the heaviest possible use. Metering by token keeps light users and heavy producers on the same plan structure without either one subsidising the other.
The free tier gives 150 fast tokens a day and stops there — enough to test the platform, not enough to run a client workflow. The three paid tiers scale the daily allowance into a monthly one: Essential at $12 a month includes 8,500 fast tokens with a 25,500 rollover cap, Premium at $30 a month includes 25,000 fast tokens with a 75,000 cap, and Ultimate at $60 a month includes 60,000 fast tokens with a 180,000 cap. Unused tokens up to that cap carry into the next month rather than resetting to zero — which matters most for anyone whose workload is seasonal rather than steady.
Two other numbers decide whether a plan actually fits a workflow, and neither shows up in the headline price. Concurrent generation limits cap how many jobs can run at once — two on Essential, three on Premium, six on Ultimate — which matters the moment a project needs a batch of variations rather than one image. And Premium and Ultimate both include 'unlimited' relaxed generation on a specific list of models, meaning image and video jobs on those models don't draw down the token balance at all, only queue a little slower when demand across the platform is high.
The honest way to choose is to work backward from a real week rather than the marketing language on each tier. Someone posting a handful of social graphics a month sits comfortably inside Essential's rollover. Someone producing daily client video drafts will burn through Essential's allowance inside the first week and spend the rest of the month buying top-up tokens at a worse rate than simply upgrading would have cost.
Dev ends up on Premium, not because it is presented as the popular choice, but because his agency runs three or four jobs at once most days, and the concurrency cap — not the token count — was the number that would have actually stopped his team.
Chapter 3
Locking a Face So It Doesn't Change Twice
Ask an AI image model for the same person twice and it will usually hand back two different faces — close in spirit, wrong in the details a viewer actually notices. For a single hero image that is a curiosity. For a comic strip, a mascot, a recurring brand spokesperson, or a storyboard with the same character in twelve panels, it ends the project before it starts.
Leonardo's answer to this is to give a face a place to live outside any one generation. The simplest version is anchoring: a reference image plus a prompt, so the model has something concrete to match rather than reinventing the person from a text description alone. A stronger, more durable version is training an Element — Leonardo's implementation of LoRA, a small model trained on a curated set of images that learns the visual pattern of a specific face, product, or style and can be called back into any future generation on demand.
The two are not really substitutes for each other. Anchoring works well within a single session and a handful of images. An Element is worth training when the same face, product, or house style needs to reappear across weeks of separate projects — a recurring mascot, a client's product line, an agency's own visual signature that has to look the same whichever team member is generating that week.
Three ready-made workflows sit on top of this system rather than requiring it to be built from scratch each time. The Consistent Character blueprint generates the same locked character in new situations with one click. Place Person In Scene changes the background around an existing character without touching their face. Multi-Character Scene Builder goes further, combining up to five separately anchored characters into a single composed scene — the tool a storyboard or a group product shot actually needs, and not something obvious from the character page's own headline copy.
None of this removes the editing work entirely. A locked face still needs its expression, wardrobe, and pose adjusted scene by scene, which is what the platform's image editor is for. What it removes is the much larger problem — starting from a blank slate every time and hoping the model remembers what it drew an hour ago.
Whether to anchor once or train an Element properly is really a question about how many more times this face needs to exist. Get that answer right early, and the rest of a multi-scene project stops being a fight.
You may also like:
Choosing the Right Luma Model Before You Waste a Generation
Chapter 4
Writing a Prompt the Model Won't Guess At
A prompt is not a request in the way a conversation is a request. These models are pattern-matchers reading a string of descriptive commands, not a listener waiting to be asked politely — which is why 'could you please make an image of a cat' produces a worse result than simply describing the cat. The polite framing is noise the model has to filter out before it gets to anything useful.
Every prompt that actually lands does three jobs at minimum: it names a subject, places that subject in a context, and states a style. Leave out the context and the model invents one — an armchair on its own becomes a generic stock-photo armchair in a generic room, because nothing told it otherwise. Add the room, the light, the materials, and the same request becomes a specific, repeatable image rather than a guess.
Three habits quietly ruin an otherwise reasonable prompt. The first is vagueness — a bare description with no setting, action or mood, which forces the model to fill every gap with whatever is statistically most common in its training data. The second is indecision — asking for a knight with a sword or an axe produces neither weapon cleanly, because the model tries to satisfy both requests at once and blends them into something usable for neither. The fix for indecision is not a better prompt; it is two separate prompts. The third is over-complexity — stacking five conflicting style references and a dozen details into one sentence overwhelms the model's ability to prioritise any of them, and the safer move is picking three to five core elements and refining from there rather than front-loading everything at once.
Style deserves its own line of thought, because a model that is not told a style will pick one anyway — usually a photorealistic default, since that is what dominates most training data. Naming a style explicitly, whether that is a broad category like 'watercolor illustration' or something narrower like 'isometric 3D render', hands that decision back to the person writing the prompt instead of leaving it to statistical default.
Aspect ratio is the one setting that does not need to live inside the prompt text at all on this platform — it is chosen from a settings menu before generation, with an advanced custom option available for anything outside the standard ratios. Worth knowing before copying prompt formats from other tools, where the ratio often does need to be typed.
Chapter 5
Turning a Still Into Motion That Holds Together
A still image only has to be convincing for the fraction of a second someone looks at it. A video has to stay convincing for every one of the next several seconds, and that is where most AI-generated motion falls apart — a face drifts slightly with every frame, a shirt changes colour mid-shot, a car turns a corner without its weight shifting onto the outer tyres the way a real car's would.
The realism a video needs breaks down into a handful of specific things a model either handles or doesn't. Skin and translucent surfaces need to scatter light internally rather than bounce it off a flat surface, or they read as plastic. A character's face and clothing need to stay locked frame to frame, or the drift becomes distracting within a few seconds. Physics — weight, gravity, the way objects should collide — needs to follow causal logic the viewer's eye already expects. And increasingly, sound needs to be generated in step with the visual, because mismatched audio breaks the illusion as fast as morphing geometry does.
No single model on the platform wins every one of these categories, which is why the practical approach is naming the specific risk in a shot before choosing a model rather than after. A close-up on a face with subtle expression needs a model built for identity consistency. A commercial with a product that must stay visually identical throughout needs a model with strong prompt adherence and native audio. A dynamic action sequence needs a model that specialises in physics and motion rather than facial nuance.
This is also the part of the platform most exposed to sudden change, and worth treating that way. Sora 2 sat on Leonardo for most of 2026 as a strong option for character consistency and synchronised dialogue, then disappeared in July after OpenAI changed its provider terms — nothing a user did, just an agreement between two companies ending. Veo 3.1 and Seedance stepped in as the recommended paths for the same kind of work, but the underlying lesson holds regardless of which model is currently favoured: a client-facing video workflow built around one specific named model is one licensing change away from needing a rebuild.
The safer habit is prototyping cheap and fast — using the lighter, quicker version of a model family to test pacing and framing — before committing tokens to a full-quality render on the heavier model. It costs almost nothing to be wrong at the draft stage and considerably more to be wrong at the final one.
Chapter 6
From Draft to Production-Size Image
A generated image that looks sharp on a laptop screen can fall apart completely the moment it needs to fill a billboard, a full-page magazine spread, or even a large digital display. Standard image upscalers, built and trained on real photographs, make this worse rather than better when pointed at AI output — they misread the model's own synthetic grain and texture as noise, and 'fix' it by smearing or distorting exactly the detail that made the image work in the first place.
Leonardo's Pro Upscaler exists specifically for that mismatch. Rather than treating an AI-generated image like a photograph, it is built to recognise the particular fingerprints of synthetic texture and clean them up instead of amplifying them, and it can push a source image up to roughly 20MB in size into a true 105-megapixel output — enough resolution to cover almost every practical digital and print use, from a full-HD display up through an outdoor billboard at readable DPI.
The upscaler splits into two working modes, and choosing the wrong one for the source image is the most common way to waste the process. Precise mode enlarges an image without changing its composition, which is the right choice when the original generation is already accurate and nothing about the identity or layout should shift. Creative mode goes further, actively repairing structural problems — a malformed hand, a distorted background face — by reconstructing that section of the image, with three strength settings running from a light touch up to an aggressive rebuild.
Inside Precise mode sits a detail easy to miss entirely: a toggle called Fix AI Image Artifacts. Left off, Precise mode acts as a pure magnifier and will scale up whatever synthetic grain or compression noise the original image had, sharpening the flaws along with everything else. Switched on, it applies a pipeline built specifically to recognise and clean that synthetic noise instead of enlarging it. The difference matters most for images generated on models known for a heavier stylised or filmic grain, and matters least for something already clean, like a real photograph or an output from a model built for crisp editing.
None of this changes the composition decisions made earlier in a project — it only decides whether the final file can survive being printed at the size the client actually asked for. A 105-megapixel file that reproduces the original grain faithfully is not automatically the right choice over one that repairs it; the decision depends entirely on what that specific image needs fixed.
You may also like:
Everything OpenArt Actually Builds Behind One Login
Chapter 7
Who Owns What You Make
Mara runs a small sticker shop and wants a straight answer to one question before she generates anything: if she makes a design on Leonardo, can she actually sell it, and can someone else legally sell the exact same thing back to her customers a week later?
The answer splits cleanly along the free-versus-paid line, and it is worth reading carefully before building a business around either side of it. On the free plan, generations are public by default, and as between the user and Leonardo, Leonardo holds ownership of those assets. A free user can still use what they generate commercially — print it, sell it, put it on a product — but because the generation is public, other users can see it and may be inspired by it or remix it for their own projects. Nothing stops a design made this way from turning up, in some altered form, on someone else's storefront.
Paid subscribers get a materially different arrangement: full ownership of their own assets either way, plus the choice to mark each piece of content private or public. Private content is visible only to the account holder and their authorised team, and Leonardo will not use it for model training or any other purpose without separate written consent — the setting to use for anything proprietary, client-facing, or exclusive. Public content, even on a paid plan, still belongs to the creator, but making it public grants Leonardo a licence to use, display and distribute that specific piece, and puts it in front of other users the same way free-tier content is.
One nuance is easy to miss in either scenario: exclusive ownership applies to the specific image generated, not to the idea behind it. Because the same prompt fed into the same model can plausibly produce a similar-looking result for someone else entirely, owning one particular generated file does not extend to blocking a separate, similar file another user creates independently.
Mara's actual decision comes down to what she is selling. A one-off concept sketch she is happy to see imitated costs her nothing to generate on the free tier. A signature product line she wants exclusively hers — the kind of asset a repeat customer would recognise as belonging to her shop specifically — needs a paid plan with the content set to private, from the first generation onward rather than after the fact.
Chapter 8
Building Leonardo Into Something Else Entirely
A founder building a print-on-demand app does not want to type prompts by hand for every customer order — they want their own product to take a customer's description and generate a design automatically, at whatever scale the business reaches. That is a different problem from the one every other chapter in this book has been solving, and it has a different door: the API.
Using a generative model manually means sitting at a keyboard, writing a prompt, and downloading one image at a time. An API removes the person from that loop entirely — a piece of code sends the request and receives the finished asset back, which is the difference between making one image for yourself and shipping a feature that generates thousands of images for other people's customers.
Most generative APIs assume the person building against them is comfortable reading dense documentation and guessing at parameters through trial and error — write code, run it, see what comes back, adjust, repeat. Leonardo's API supports that code-first path fully, but it also supports a second one: designing the exact result visually inside the normal app interface first, then copying the working configuration straight into code once it looks right. That second path removes the guessing entirely for someone who knows what they want the output to look like but isn't an engineer by trade.
The businesses actually building on this pattern cluster around a few repeatable shapes. Print-on-demand platforms let a customer describe a design in their own words and generate a finished mockup on a t-shirt or mug before purchase, which measurably increases conversion because the customer can see the exact product rather than imagining it. Clothing retailers use the same underlying capability to show how a garment looks across different body types and skin tones without a separate photoshoot for each. Property platforms generate furnished versions of empty rooms directly from a listing photo, letting a buyer picture the space instead of touring it bare. At the largest end, Coca-Cola's own real-time holiday campaign ran on a custom model built on this same underlying engine, generating a personalised digital snow globe for every user's conversation rather than serving one static asset to everyone.
The founder deciding between these two paths does not need to pick one permanently. Prototyping the exact visual result inside the app, confirming it looks right, and only then converting it into a production API call is a legitimate workflow on its own — not a beginner's shortcut to be abandoned once the business grows, but the actual recommended order of operations either way.
Questions readers actually ask
Can I sell images I make on Leonardo's free plan?
Yes — free-tier generations can be used commercially, but they are public by default and, as between you and Leonardo, Leonardo owns the asset. Because it's public, other users can see and potentially reuse a similar design.
Does upgrading to a paid plan make my work exclusive?
Only if you also set it to private. Paid subscribers own their assets either way, but public content on a paid plan still grants Leonardo a licence to use and display it, and other users can see it. Private content stays visible only to you and isn't used for training without separate consent.
What is the difference between Precise and Creative upscaling?
Precise mode enlarges an image without altering its composition — the choice when the source is already accurate. Creative mode actively repairs structural problems, like a distorted hand or background face, at a chosen strength from subtle to strong.
How large can an upscaled image actually get?
Up to a true 105 megapixels, starting from a source file of roughly 20MB or smaller — enough resolution for most print formats up to large outdoor billboards at usable DPI.
Do unused tokens carry over to the next month?
Yes, up to a cap. Essential's rollover ceiling is 25,500 tokens, Premium's is 75,000, and Ultimate's is 180,000 — beyond that ceiling, additional unused tokens do not accumulate further.
How many characters can appear together in one locked scene?
The Multi-Character Scene Builder blueprint combines up to five separately anchored characters into a single composed scene.
What happened to Sora on Leonardo?
Sora 2 and Sora 2 Pro were removed from the platform in July 2026 following a provider-side change at OpenAI, not a decision by Leonardo. Veo 3.1 and Seedance were positioned as the direct replacements for that kind of work.
What does 'unlimited relaxed generation' actually include?
Only a specific list of models, and only on the Premium and Ultimate plans — it doesn't draw down the token balance, but generation can slow during high platform demand to keep access fair across all users.
Is there a way to use Leonardo without the app interface at all?
Yes, through the API, which supports both a code-first workflow using standard documentation and endpoints, and a visual-first workflow where the exact result is designed inside the app first and the working configuration is then copied into code.
Do I need to buy a paid plan to train a custom Element?
The pricing page ties a specific number of personal AI models — trained Elements — to each plan tier, from 10 on Essential up to 50 on Ultimate, so training one requires at least the entry paid tier.
Contact / More useful information from RamthaMedia
Official source links:
Leonardo
The token allowances and plan prices in this book reflect what Leonardo published at the time of writing. Token limits and pricing tiers are exactly the kind of thing a platform revises without much notice — check the current pricing page linked below before choosing a plan.
As an Amazon Associate, RamthaMedia earns from qualifying purchases.
Disclaimer: This eBook is compiled from publicly available information and was accurate at the time of writing. For full and up-to-date details, please visit the official website linked above. RamthaMedia accepts no legal liability for any decision made on the basis of this eBook, and nothing here is professional, financial or legal advice. The image used for the cover page is illustrative only – a stock photo from Pexels or an AI-generated image, never a real photograph of the site described.