Ideogram, From First Prompt to Finished Product Photo

See what Ideogram's background remover, character tool, custom models and API do before you build with it.

By RamthaMedia

RamthaMedia Free eBooks  ·  August 2026

Price: Priceless
 ·  16 min read

Preface

A product photo needs its background gone before a listing goes live tonight. A campaign needs resizing for six ad placements before launch. A brand needs its exact colors to survive a hundred new designs without one off-brand shot slipping through. This book walks through what Ideogram actually generates, edits, and automates — the background remover, character locking, custom brand models, print-ready transparency, and the API underneath all of it — so the next asset gets made today, not scheduled for a shoot next month.

Chapter 1

A Product Photo That Has to Lose Its Background Tonight

It's eleven at night and a ceramic mug is sitting on a kitchen counter, propped against a stack of mail, waiting to become a marketplace listing before tomorrow's featured-collection deadline closes. There's no lightbox, no white sweep, and no time to book one.

Most background removers handle this by classifying every pixel as either subject or background and cutting along that line. The problem sits exactly at the edge, where the two blend — a strand of hair scattering light, a handle catching a reflection, a thin ceramic rim. A segmentation model has no real understanding of what it's looking at, so it guesses at the border and the guess shows up as a halo or a smear of leftover color.

Ideogram approaches this differently: it's a generative model trained to reconstruct the border itself, pairing each edge pixel with a level of transparency and the correct foreground color underneath it, rather than deciding pixel-by-pixel whether something belongs in or out. The mug's rim, the reflection in the glaze, even a stray hair caught in the frame — all of it comes out clean because the model has some sense of what a ceramic surface or a strand of hair is supposed to look like.

In the browser, this is free — upload the photo, get back a transparent PNG. Through the API, it's a single POST to /v1/remove-background at a stated $0.01 per image, returning a signed link to the result that stays valid for roughly a day, long enough to download and store it.

The site's own published comparison runs the same source photo through several well-known removal tools side by side, and calls out glass, reflections, and fine typography as where the difference shows most; that's Ideogram's own claim about its own benchmark, worth treating as a starting point rather than a verdict.

The mug is sorted. But a seller with two hundred SKUs doesn't want to do this one photo at a time — and a business that needs the same model, mascot, or spokesperson to show up consistently across forty different scenes is facing a harder version of the same problem: not just a clean edge, but a face that has to stay the same face every time.

What you can actually do here

Ideogram is really several tools sharing one login. This is what each one is actually for, and where to find it.

Cleaning up images for selling

Use Who it fits Where Worth knowing
Remove a product photo's background Marketplace sellers, e-commerce catalog managers Apps → Background Remover → upload → download PNG Preserves hair, glass and fine edges without a halo
Free in the browser; $0.01 per image through the API
Erase a photobomber or watermark from a photo Photographers, stock cleanup, listing photos Apps → Object Remover → brush the object → generate Rebuilds shadows and texture behind the removed object
Averages under ten seconds per removal in the site's own benchmark

Keeping one face or one look everywhere

Use Who it fits Where Worth knowing
Lock a character's face across dozens of images Indie game devs, brand mascots, storybook illustrators Character tab → upload one reference photo → prompt Needs one photo, not a training set
Complex poses or low-quality references can still drift
Train a model on a brand's exact visual identity In-house creative teams, agencies with recurring clients Models tab → new model → upload 15–100 images → train Output comes back on-brand without re-prompting every detail
Self-serve tops out at 100 images; bigger datasets need Enterprise

Getting a design out the door

Use Who it fits Where Worth knowing
Resize one ad creative into every placement it needs Performance marketers, agencies running multi-channel campaigns Apps → Ad Resizer → upload → choose placements → export Rebuilds the layout instead of cropping or stretching it
Priced on token usage, not a flat per-image fee
Generate print-ready art with a transparent background Print-on-demand sellers, packaging and merch designers Generate → end the prompt with "transparent background" → upscale to 8K Alpha channel survives the upscale
Leaving out the exact phrase returns a colored backdrop
Fix a typo or swap a headline without regenerating the design Poster, book cover and social graphic designers Generate a design → Layerize Text → click a line → retype The rest of the composition stays untouched
Still in beta; heavily stylized or curved lettering may not be detected

Building on top of it

Use Who it fits Where Worth knowing
Call image generation from your own product Developers building generation into an app or backend developer.ideogram.ai → POST /v1/ideogram-v3/generate One endpoint covers generate, remix, edit, and reframe
Priced per model and rendering tier chosen, not a flat credit rate
Let an AI assistant operate Ideogram directly Anyone already working inside Claude, ChatGPT, or Cursor Connect the MCP server at https://mcp.ideogram.ai/mcp The assistant chains tools without anyone writing integration code
Draws from the same subscription credits as the web app

Chapter 2

The Same Face, Forty Times

Generate the same person twice from the same prompt and two different faces come back. For a storybook, a game, or a brand mascot that has to appear across dozens of scenes, that's not a stylistic quirk — it ends any story that depends on one recognizable character.

The usual fix is training a small custom model per character: gather ten or more reference photos, fine-tune, wait. Ideogram Character skips that step. Upload a single reference photo, write a prompt describing a pose, outfit, or scene, and the same face carries through the output. According to independent creator tests the site cites, run across pose changes, expression changes, outfit swaps, and a stress-test cinematic prompt, Ideogram was the only model in that comparison that held the character's identity through every round rather than drifting partway.

The mechanism sits on top of the same generation pipeline: a Character API endpoint accepts a character_reference image URL alongside the usual prompt, so the identity lock works the same way whether it's triggered from the web app or called from code. That makes it usable for a run of individual LinkedIn-style headshots, a set of branded mascot posts, or a storybook's illustrations, all from one uploaded photo.

Two related tools extend it. Face Swap drops the locked character into an existing background or scene using Magic Fill, useful when the setting already exists and only the person needs to change. Remix borrows the composition and mood of a reference image and reapplies it to the character, which is a faster route to a consistent look than describing lighting and framing from scratch each time.

There are limits worth knowing before relying on it for a shoot-replacement job: a low-quality or heavily angled reference photo can still cause drift, and fine control over exactly how the character's hair, clothing, or accessories render requires manually adjusting a character mask rather than trusting the default. It works from one photo, not from a trained dataset — which is also why it can't yet guarantee pixel-perfect wardrobe consistency the way a dedicated custom model can.

A locked face solves identity. It says nothing about the words sitting inside the image next to that face — the sign behind the mascot, the headline over the headshot — and that's a separate problem Ideogram treats as its own specialty.

Chapter 3

Text That Holds Its Shape

A poster comes back from generation with almost everything right — the composition, the lighting, the mood — except the one line of text across the top, which renders as a garbled approximation of real letters. The usual response is to regenerate from scratch and hope the next pass gets luckier.

Most image models struggle with in-image text because they were trained on loose, descriptive captions that never had to specify exactly what a rendered word should say. Ideogram 4.0 was trained differently: on structured JSON captions that label a scene as organized, tagged data rather than a single sentence — a high-level description, a style block, and a list of individual elements, each with its own description and, where relevant, a literal string of text to render. That structure is also why the model is described as more literal than earlier versions: vague words like "moody" or "editorial feel" carry less weight now, and specific direction — a named light source, an exact hex color — gets followed more reliably.

Plain-text prompts still work; they pass through a translation layer called Magic Prompt that expands a normal sentence into the JSON structure automatically. Sending a well-formed JSON prompt directly bypasses that translation, so what's written is what renders, without the model reinterpreting intent. This matters most exactly where text is involved: quoting a headline in double quotes tells the model to treat it as a literal string rather than a suggestion.

Once text is in the image, editing it used to mean regenerating the whole design. Layerize Text changes that: pressing it turns each line of rendered text into its own selectable, editable component, while the rest of the visual design stays exactly as generated. A headline can be reworded, a font swapped, a size adjusted, without touching anything else on the poster.

It's available on every plan including the free tier, and through the API for anyone building a design tool on top of it. Its real limit is that it's still in beta: it works reliably on clear, straight typography and can miss curved, heavily stylized, or graphic-embedded lettering, in which case the honest fallback is to keep that particular line as a raster image rather than force a layer extraction that won't detect it.

Getting a design's words right is one problem. Getting the same design correct across six different aspect ratios — a square Instagram post, a tall Story, a wide banner — without redrawing it each time is the next one.

Chapter 4

One Ad, Every Shape It Needs to Be

A campaign creative is finished and approved, and now it needs to exist as a leaderboard banner, a mobile banner, a skyscraper, an Instagram Story, and a landscape in-stream video frame — five different aspect ratios, all supposed to look like the same campaign.

The naive approach is cropping or stretching, and both fail visibly: cropping cuts off the logo or the subject, stretching distorts everything. Ad Resizer takes a different approach — it reads where the subject, the logo, and the copy sit in the source creative, then rebuilds the layout for the target shape and extends the background to fill whatever space the new frame adds, rather than cropping or stretching the original.

The workflow is upload one creative in whatever shape it started in, pick from the supported preset placements — the standard IAB display sizes plus Instagram feed, Stories, Reels, Facebook feed ads, YouTube, connected TV, and print — and export the full set at once. Color, font, and key visual elements are held across every size in the set, so the ad reads as one campaign rather than several separately edited versions drifting apart.

Pricing runs on token usage rather than a flat per-image charge: the total combines the cost of reading the source creative's text and image content with the cost of generating each resized output, and every additional variation adds to that image-output cost. It's not a subscription line item so much as a per-job calculation, which matters for anyone estimating cost across a large placement list rather than one or two sizes.

This solves distribution — the same message, correctly shaped for wherever it runs. It doesn't solve the separate problem of getting a design ready for physical production, where the requirements shift from aspect ratio to resolution, transparency, and print-safe detail.

You may also like:
Turning One Song Into Ten Videos With Kaiber

Chapter 5

From Blank Shirt to Storefront Listing

A print-on-demand seller wants a new design live today — not next week, and not after three rounds of a freelancer's revisions. The bottleneck isn't the idea; it's turning that idea into a file a printer and a listing template will both accept.

Ideogram generates transparency natively rather than as a post-processing step: adding the phrase "transparent background" to a prompt produces a clean alpha channel directly, with no separate background-removal pass and no haloing on thin strokes or semi-transparent elements. That single phrase is required exactly as written — leaving it out returns a colored backdrop instead.

Editing preserves that transparency. Describing a change — swap the colorway, add an element, adjust a detail a customer requested — applies through the same transparent file without needing to remove the background again afterward, which turns one design into a small collection of variants rather than a one-off.

For print, standard t-shirt printing needs upward of 600 DPI, which most generated images don't start at. Upscaling to 8K clears that bar for posters and all-over prints as well as smaller formats, and the transparency carries through the upscale rather than being lost in the process.

Once the design exists, Ideogram Edit can place it directly onto a product: describing the scene — printing the design onto a white t-shirt and generating a photo of the finished shirt — produces a mockup without a Photoshop template or a separate mockup generator, useful for previewing a listing before committing to it.

The site's own guidance on prompt structure is worth following closely here, because POD designs share a specific shape that performs better than a pile of adjectives: lead with a subject doing something, land a deadpan twist, name three to five flat colors up front rather than describing a mood, close with one phrase naming the style, and end every prompt with "transparent background." Committing to one register — illustration or photoreal, never both in one prompt — avoids the muddy middle ground that mixing produces.

A design that's finished and print-ready still has to be paid for through one of Ideogram's plans, and that's where the credit system, not the tools, becomes the thing worth understanding before scaling up.

Chapter 6

What a Credit Actually Buys

Someone signs up, sees "unlimited slow credits" on the free tier and "priority credits" on every paid plan, and has no immediate way to tell what either of those actually means for a day of real work.

The free tier includes 10 slow credits a week and one concurrent generation — enough to explore the tool, not enough for production work with a deadline. Plus is priced at $20 a month, or $15 a month billed annually, and includes 1,000 priority credits monthly, private generation, unlimited slow credits, unlimited character consistency, and quality export. Pro runs $60 a month ($42 annual) with 3,500 priority credits, batch generation, and the largest generation queue. Team is priced per user at $30 a month ($20 annual) with 1,500 priority credits per user, plus central billing and early access to collaboration features. Enterprise is a custom, post-paid arrangement with a negotiated credit amount, private custom models, and volume discounts on the API.

The practical difference between slow and priority credits is speed and reliability: slow credits queue behind everyone else's slow generations, while priority credits jump the line. How many credits a single generation costs also depends on which model and rendering setting is chosen — a 4.0 image at Turbo rendering runs 2 credits per image, while Quality rendering on the same model runs 6 credits per image, and older models like 1.0 cost less per generation across the board.

Top-ups exist for anyone who runs out mid-month rather than wanting to upgrade the whole plan: Plus offers 150 priority credits for $4, Pro offers 250 for the same $4, and Team offers 250 per top-up as well.

One number worth checking directly on the pricing page rather than assuming: whether unused priority credits carry over month to month is listed as one of the site's own frequently asked questions, but the answer text itself wasn't part of what this book could confirm — it's the kind of detail to verify before budgeting a month of heavy production around a leftover balance.

Credits pay for generation. They say nothing about what happens when a brand needs every generation, forever, to look like it came from the same design team — which is a training problem, not a pricing one.

You may also like:
What Google Flow Actually Lets You Build

Chapter 7

Training a Model That Knows Your Brand

A creative lead at a mid-size brand keeps getting outputs from general-purpose AI tools that are technically fine and instantly recognizable as generic — the palette, the composition, and the finish all belong to no one in particular. Feeding the brand's actual product photography and design system into a general model doesn't fix this, because a model trained on the whole internet has no reason to prefer one brand's aesthetic over the millions of others it also learned from.

Custom Model Training addresses this directly: fine-tuning a model exclusively on a brand's own assets so its output default matches that brand's style rather than needing to be steered toward it with every prompt. The self-serve process is upload 15 to 100 images through the Models tab, let the system auto-generate captions, review and refine those captions, and start training. Quality of the training images matters more than sheer volume — the site is explicit that forty well-captioned images will outperform a much larger set of mediocre ones.

The same process runs through the API for teams that want it scripted: create a dataset, upload the images (a minimum of ten, a maximum of one hundred, with optional caption files matched by filename), start training with a model name, and poll the model's status until it reads COMPLETED. Once training finishes, that model's custom_model_uri is passed into the same generate endpoint used for ordinary prompts.

On the Team plan, a trained model can be shared across an organization rather than owned by one person — anyone can train a model, share it with the group, and the whole team generates from the same brand foundation without a bottleneck.

Enterprise sits above self-serve for anything larger or more sensitive: datasets in the thousands of images, custom captioning pipelines, layerized and SVG rendering of exact logos and typography, and a stated commitment that training data and the resulting model belong to the customer rather than feeding into Ideogram's shared models. Enterprise data handling is covered under SOC 2, a Data Processing Agreement is available, and data retention is configurable down to as little as 30 minutes for enterprise customers who need it.

None of this training happens through the browser alone if a team wants it wired into an existing pipeline rather than a manual upload — and that's where the API, and the newer option of an AI agent operating the tools directly, becomes the more relevant chapter.

Chapter 8

Calling Ideogram From Code, or Just Asking Claude

A developer is building a storefront that needs to clean and generate thousands of product images automatically, overnight, without a person clicking through the web app once per image. That's a job for the API, not the browser.

The core generation endpoint is a single POST to /v1/ideogram-v3/generate, authenticated with an Api-Key header, accepting a prompt and optional parameters like a character_reference or a custom_model_uri. Alongside it sit dedicated endpoints for the more specialized jobs covered earlier: /v1/remove-background for a transparent PNG at $0.01 per image, /v1/remove-object for masked object removal, and the dataset and training endpoints for custom models. Every response carries an x-request-id header, which is worth logging on every call — it's what turns a vague support ticket into a specific, investigable failure.

A few practical constraints show up across these endpoints: uploaded images are capped at 10 MB for the background remover, JPEG, PNG, and WebP are the accepted formats, and returned signed URLs are valid for roughly 24 hours, so downloading the result promptly rather than hotlinking it is the safer pattern for production use. Running requests in parallel is the documented approach for processing a full catalog, with 429 rate-limit responses treated as retryable with backoff rather than a hard stop.

The newer option is the Model Context Protocol server, which lets an AI assistant operate Ideogram's tools directly inside a conversation rather than through a script someone wrote. Connecting it is pointing a compatible client — Claude, Claude Code, ChatGPT, Cursor, Cline, or OpenCode — at https://mcp.ideogram.ai/mcp; the first connection opens a browser window for an OAuth sign-in, and after that the assistant has access to generation, editing, and collection tools it can chain together inside one conversation. A single brief like producing an asset pack — a hero image, several social cards, and print posters — can come back from one exchange rather than a sequence of separate manual steps.

Usage through MCP draws from the same subscription credits as the web app; there's no separate billing tier for going through an assistant instead of a script. Which route makes sense depends on the job: MCP suits a person directing an agent through their own account, while the REST API is the better fit for a server-to-server pipeline that needs to run unattended under its own credentials.

Whichever route generates the image, the output still belongs to someone, and what that someone is and isn't allowed to do with it is governed by a separate, and easy to skip past, set of policies.

Chapter 9

What You're Allowed to Do With What You Make

Someone finishes a batch of AI-generated product photography for a paying client and pauses on a question that's easy to assume the answer to: who actually owns these images, and is there anything about how they were made that the client needs to be told?

Ideogram doesn't claim ownership over what a user generates. Output can be used commercially without restriction tied to the generation itself, and the site's own terms assign any rights it might otherwise hold in that output back to the user. The one boundary worth knowing sits on the input side and the training side, not the output: using generated output to train or build a competing image model is explicitly against the usage policy, and the site encourages — without strictly requiring in every case — disclosing that an image was AI-generated if it's being distributed in a way that could otherwise mislead someone about its origin.

The usage policy also rules out a specific list of misuses regardless of plan: content that's illegal, harmful to children, non-consensual imagery of real people, or material designed to mislead people about whether something is a real event. Scraping the service with bots, reverse-engineering the underlying technology, and building a competing product on top of its outputs are also explicitly restricted.

Above the output-ownership question sits a separate one for anyone wanting to run Ideogram's actual model weights on their own infrastructure rather than through the hosted app: that's governed by licensing, not the usage policy, and it comes in three tiers. A Non-Commercial license is free and covers research, evaluation, and personal projects using the public, quantized weights on Hugging Face — no commercial use permitted under it. A Self-Serve Commercial License adds the right to self-host those same public, quantized weights for commercial use, capped at 100,000 generated images a month, sold as a monthly image allowance selected at checkout. Anything larger — full-precision weights, customer-facing products built on top of the model, reselling access to third parties, or custom legal terms — moves into Enterprise, which is negotiated directly rather than purchased through a checkout flow.

For nearly everyone reading this book — a seller cleaning up photos, a designer resizing a campaign, a brand training a model on its own catalog — none of the self-hosting licensing tiers come into play at all; using the hosted app or the hosted API already carries commercial rights to what gets generated. The licensing tiers only matter the moment the plan shifts from using Ideogram's servers to running the model on someone else's — at which point it's worth reading the specific tier's terms rather than assuming the hosted app's rules carry over unchanged.

Questions readers actually ask

Is Ideogram's background remover actually free?

Yes, for the browser tool — it's free to use on the website with no subscription required. Volume use through the API is priced per image, currently stated as $0.01 per image.

How many reference photos does Character need?

One. Ideogram Character is built to work from a single reference photo rather than a multi-image training set or a fine-tuned model.

Can I use Ideogram-generated images commercially?

Yes. Ideogram doesn't claim ownership of generated output, and using it commercially is permitted under the usage policy, with the exception of using output to train a competing image model.

Does using the MCP server cost anything extra?

No — usage through the MCP server draws from the same subscription credits as the web app. There's no separate billing for requests made through an AI assistant.

What image formats does the background and object remover accept?

JPEG, PNG, and WebP, up to 10 MB per image for the background remover.

How many images do I need to train a custom model?

The self-serve tool accepts 15 to 100 images. The site notes that image quality and captioning matter more than the total count — a smaller set of well-captioned images can outperform a larger, weaker one.

Is the Ideogram 4.0 model actually open source?

The quantized weights are published on Hugging Face under a free Non-Commercial Model Agreement for research and personal use. Commercial use of those same weights requires either a Self-Serve Commercial License or an Enterprise agreement.

Why does Ideogram want prompts written as JSON?

Ideogram 4.0 was trained on structured JSON captions rather than plain sentences, so a JSON-structured prompt is used exactly as written, while a plain-text prompt is first expanded automatically by a translation layer called Magic Prompt.

Can I edit text in a design after it's generated?

Yes, using Layerize Text, which turns each line of rendered text into a separate, editable component without regenerating the rest of the design. It's currently in beta and works most reliably on clear, straight typography.

Contact / More useful information from RamthaMedia

    Official source links:
    Ideogram

    The plan and API prices in this book were accurate on Ideogram's own pricing page when this book was written. Subscription tiers, credit allowances, and per-image API rates change from time to time — check the official pricing page linked below for the current figures before you commit to a plan.


    Disclaimer: This eBook is compiled from publicly available information and was accurate at the time of writing. For full and up-to-date details, please visit the official website linked above. RamthaMedia accepts no legal liability for any decision made on the basis of this eBook, and nothing here is professional, financial or legal advice. The image used for the cover page is illustrative only – a stock photo from Pexels or an AI-generated image, never a real photograph of the site described.

    RamthaMedia
    RamthaMedia

    About the Founder – A. Ravinder
    A. Ravinder is the Founder, Author, Digital Publisher, and Editor-in-Chief of RamthaMedia, a Telugu-focused digital media and publishing platform dedicated to delivering trusted news, practical knowledge, books, and smart buying guides.
    With strong experience in digital publishing, journalism, content research, and affiliate product analysis, he creates reliable, easy-to-understand, and value-driven content that helps readers make informed decisions in their daily lives.
    Through RamthaMedia, he combines news reporting, book publishing, educational resources, and honest product reviews — building a trusted knowledge ecosystem for Telugu and Indian audiences.

    Articles: 302