Create with GPT Image 2 from text or reference images

GPT Image 2 is an OpenAI image generation model for creating a new image from a prompt or working from existing visual material. On aiimg.me, choose the workflow that matches your starting point, review the ratio, output tier, result count, and credit estimate, then generate in the browser.

Decorative collage representing text-to-image and reference-led creation

Start with the input that already holds the clearest direction

One model for new images and reference-led edits

OpenAI documents GPT Image 2 as accepting text input and image input and output, with image-generation and image-edit operations. aiimg.me presents those capabilities as two focused browser workspaces: text-to-image and image-to-image. The settings on this page describe aiimg.me, not every option in the official API.

The model is useful while the visual is still being decided: the subject, setting, composition, palette, material, or treatment can all be described or supplied as reference context. It is a poor substitute for deterministic editing when a finished file must remain unchanged outside one exact correction.

Read the official GPT Image 2 model documentation

Useful when you need to

  • Turn a product, campaign, editorial, or social brief into a first visual direction.
  • Explore alternative settings, palettes, materials, framing, or lighting before production.
  • Use a product shot, sketch, previous result, pose, or layout as context for a new version.
  • Compare a small set of candidates in the aspect ratio required by the final placement.

Use another tool when you need

  • Layers, brush masks, locked regions, or repeatable pixel-level corrections.
  • Lossless restoration or guaranteed enlargement of an existing file.
  • Exact typography, logos, measurements, or compliance-critical geometry.
  • The same character, face, brand element, or source detail preserved identically across outputs.

Text-to-image or image-to-image?

Choose the mode by asking where the most reliable information already lives. Use text when the brief can carry the direction. Use image-to-image when showing the subject, structure, or look is clearer than describing it from memory.

Use text-to-image when the brief is the source of truth

Begin with the subject and intended placement, then add the setting, composition, light, material, color, and mood. A useful prompt makes visible decisions: “Editorial product photograph of a matte cobalt desk lamp on pale limestone, soft side light, 4:5 composition, generous empty space above the lamp for a headline.”

Use image-to-image when the source carries information you need

Upload at least one image, assign each reference a role, and describe the requested change. For example: “Use reference 1 for the bottle silhouette and front camera angle. Use reference 2 only for the sandstone palette. Place the bottle on pale travertine with warm morning side light.” References guide a new generation; they do not lock every source pixel.

Choose the smallest run that answers your next question

The workspace calculates the total from the selected output tier and result count before you press Generate. One 1K result is the lowest-cost way to test whether the composition and instruction are moving in the right direction. Increase the tier or count only when that choice serves a specific next step.

Images1K2K4K
110 credits20 credits40 credits
220 credits40 credits80 credits
440 credits80 credits160 credits

1K: test the brief before scaling the run

At 10 credits per image, 1K is the economical first pass for checking the subject, framing, palette, and prompt logic. Treat it as a decision tool rather than a lesser version of an unresolved idea.

2K: request a larger working result

At 20 credits per image, 2K makes sense after the direction is established and a larger output would help the next stage of the work. It should answer an output need, not replace another prompt revision.

4K: use the highest tier only when size matters

At 40 credits per image, 4K costs four times as much as 1K. It changes the output tier; it does not correct a vague instruction, an unwanted object, incorrect text, or weak composition.

Generate one to validate; use two or four to compare

Request one image while diagnosing a prompt. Choose two or four when simultaneous alternatives will make the decision easier, knowing that the per-image cost is multiplied by the result count. The server charges each task after Generate; a task recorded as failed has its consumed credits refunded.

Set the ratio for the intended placement

Available choices are Auto, 1:1, 5:4, 9:16, 21:9, 16:9, 4:3, 3:2, 4:5, 3:4, 2:3. Use Auto while the layout is open, 1:1 for square placements, 4:5 or 3:4 for portrait layouts, 9:16 for vertical screens, and 16:9 or 21:9 for wide compositions. Pick the destination before generating when possible; a later crop can remove the space the composition depended on.

Use reference images only when each one has a job

Image-to-image requires at least one reference and accepts up to 16. Each file can be up to 30 MB; supported formats are JPG, JPEG, PNG, WEBP. The upper limit is useful when separate images supply separate facts, not as a target for every task.

Start with the smallest set that can explain the job. If two references disagree about shape, camera angle, palette, or light, state which image controls each decision. Extra inputs may add useful context, but they may also introduce competing directions.

Upload only material you are allowed to use, and review our privacy policy before adding confidential or sensitive images.

From brief to reviewed result in four steps

A useful run starts before Generate. Decide what information is reliable, choose the matching workflow, spend only enough credits to answer the next question, and inspect the output before continuing.

  1. 01

    Choose the source of truth

    Use text-to-image when the brief can define the scene. Use image-to-image when a source image carries subject, structure, pose, framing, palette, or material information that the result should draw from.

  2. 02

    Write visible decisions

    Describe the subject, placement, composition, light, materials, and mood. For image-to-image, assign every reference a job and state both the anchor and the requested change.

  3. 03

    Set the destination and budget

    Choose the aspect ratio for the intended placement, then select 1K, 2K, or 4K and one, two, or four results. Review the displayed total before submitting the task.

  4. 04

    Generate, inspect, and change one thing

    Check the full-size result against the brief and references. If another pass is needed, change the instruction that caused the largest problem instead of rewriting every variable at once.

Review the output before it becomes an asset

OpenAI's image-generation guide notes that precise text placement and clarity can still fail, and that recurring characters or brand elements may not remain consistent across generations. Those limits matter even when the overall image looks convincing. Inspect the full-size result against the brief and source material.

  • Check text, logos, labels, numbers, and symbols character by character; add exact typography later in a layout tool when accuracy matters.
  • Compare faces, characters, products, and brand elements across results instead of assuming identity or structure stayed fixed.
  • Inspect hands, patterns, reflections, symmetry, joins, dimensions, and small structural details before publication.
  • Do not expect 4K, more results, or more references to resolve an ambiguous or contradictory instruction.

Use a conventional editor for masks, locked regions, exact typography, measured geometry, lossless restoration, and final compliance-sensitive corrections.

What to try before you generate again

Why did it miss part of my prompt?

Put the must-have subject, action, and composition first. Then remove decorative phrases that do not change the picture. If the prompt asks for several people, props, camera choices, and styles at once, split the job into smaller attempts. Change one thing between runs so you can tell which instruction actually moved the result.

Why did my uploaded image change more than I asked?

Image-to-image makes a new version rather than changing a few fixed pixels. Say what must stay, then describe one main change. A line such as “keep the bottle shape and label placement; replace only the background” is clearer than a long style brief. If part of the original must remain exact, finish that part in an image editor.

Will adding more reference images improve the result?

Start with one strong reference. Add another only when it brings something distinct, such as the pose, color palette, lighting, or material. Name that role in the prompt so the images are easier to read together. The workspace accepts up to 16 references, but using the full allowance usually makes the direction harder to control, not easier.

What should I do when the text or logo comes out wrong?

Keep generated text short, put the exact wording in quotation marks, and remove instructions that compete for attention. For a logo, label, price, or headline that must be correct, use the generated image for the visual direction and add the final lettering in a layout tool. Moving to a larger resolution gives you more pixels, not better spelling.

When is 2K or 4K worth the extra credits?

Use 1K at 10 credits per image while you are still changing the subject, framing, or prompt. Choose 2K at 20 credits once the direction works and you need a larger file. Save 4K at 40 credits for a real print, crop, or large-display requirement. If the idea is wrong, revise it at 1K before spending more.

Start with a prompt, or bring the image that already matters

Open text-to-image when the brief contains the useful direction. Open image-to-image when a product, subject, composition, pose, or visual treatment is easier to show than describe. Both links select GPT Image 2 without starting a generation.