The name of the model does not tell you what it can be used for
The "latest image models" are not lined up on one shelf. There is a mixture of closed models offered through a provider called through APIs, models that can be used in the product UI, models in which only the inference code is made public, models in whose weights are distributed, and models that have additional conditions for commercial use. If we collectively refer to this as OSS, important questions such as whether it can be run locally, whether input is retained, and whether it can be distributed to teams disappear.
OpenAI announcement obtained on 2026-10-04 provides gpt-image-1 as an API and explains editing, instruction following, and character drawing as its main uses. The announcement also touches on C2PA metadata and security measures. This is a description of the API provision, not an announcement of weight disclosure or local inference. Prices vary depending on input/output tokens, quality, and size, and the latest prices and usage qualifications will be confirmed in the official price list and API document immediately before implementation.
- 1Product UI/API Usage
- 2Providers manage inference bases and updates
- 1open weights
- 2Consider inference possibilities on your own GPU/cloud
- 1OSS code
- 2Verify code access
- 1Commercial use
- 2Check weights, outputs, inputs, and local terms individually
Practical implications of referencing and editing
OpenAI's Image Prompting Document explains editing with reference images and masks in GPT Image 1, and suggests that a high input_fidelity setting aims to preserve input detail while using more image input tokens. Here, "high fidelity" does not guarantee exact match. If the reference person appears at a different angle, with different lighting, and with different occlusions, the model estimates what is not visible. For elements that cannot tolerate errors, such as the colors and logos in product photos, the process of compositing them on a separate layer is left.
Google DeepMind's model cards list lists Imagen 4 as a generated model and indicates the update date. The model card is an entry point to read the intended use, evaluation, and known limitations, and is not a document that fixes the availability areas and prices for each service. Even if the official benchmarks are good, you will need to evaluate your own proprietary products, non-English characters, fine lines, and print resolution separately.
Positioning of Stable Diffusion 3.5
Stability AI's official announcement introduced SD 3.5 Large, Large Turbo, and Medium in October 2024. Turbo is aimed at fast generation with fewer steps, and the company explains its availability under a Community License. What is useful here is the possibility of increasing the number of trials on your own PC or in a closed environment. On the other hand, do not just quote the license name and conclude that it is ``unconditionally commercially available.'' Check the current model-weight distribution page, applicable conditions such as revenue and organization size, derived models, and third-party LoRA licenses separately.
Fast local generation has operational costs that include GPU memory, drivers, quantization, resolution, batching, and even upscaling. Cloud APIs have output differences due to latency, unit price, rate limits, data handling, and model updates. Choose based on the amount of iteration, confidentiality, reproducibility, and your team's maintenance ability, not on which is better.
Create the evaluation table first
The same ten questions are asked of candidate models. Character consistency, product labels, short Japanese and English characters, hand contact, composition, reference editing, mask editing, transparent backgrounds, latency, and explainability in case of failure. Record input, output, settings, scoring, license URL, and implementation date for each problem. Although the provider's human preference benchmark serves as a basis for ``candidate'', it is not the exact condition for your own success.
- 1Use one sentence
- 2Define passing image
- 3Compare candidates with the same input
- 1Score Quality, Cost, Time and Rights
- 2Small Implementation
- 3Re-evaluate regularly
- 1Detect Model Updates
- 2View Differences in Fixed Test Sets
Verification limit
This note is a compilation of primary materials that could be viewed on the date of acquisition, and is not an implementation verification of all regions, all plans, and all SDKs. Product names and availability will be updated. At the time of implementation, review the official documentation for current model names, terms of use, pricing, data retention, output rights, and content policies, and in particular, follow the organization's consent and protection policy for human images and customer data.
Check licences again when a product changes
This note separates a hosted model from weights and a licence because that distinction remained material after 2024. OpenAI's September 2026 ChatGPT Images 2.5 announcement is an observed product update, not evidence that the API, weights, or commercial terms in this note changed. Before use, retrieve the current product-specific terms and account contract, then record the exact endpoint or UI, output rights statement, retention setting, and date.
MENTAL MODEL / REASONING ORDER
Change one thing at a time.
Decide what the image communicates, its medium and dimensions, and the meaning you want to preserve.
SOURCES
01YOUR NOTES