Can GPT-4.1 analyze images directly?
It can perform image and text understanding. When using Chat Completions, combine text and image_url content blocks in the same message, and specify the charts, areas, or issues to inspect. Results are returned as text analysis; if you need to generate or edit images, choose a dedicated image model.
Does 1 million tokens of context mean it can generate an equally long answer?
No. Context capacity and the per-response output limit are different metrics: GPT-4.1 has a native context window of up to 1 million tokens, with a maximum output of 32,768 tokens. Long-form deliverables can be generated by chapter, with space reserved for input materials, conversation history, and responses.
Should I use Responses or Chat Completions?
Applications that already use message arrays can use Chat Completions, submitting model and messages and reading the response from choices. Responses uses model and input, and handles output according to response events. The data structures of the two interfaces differ, so request bodies or parsing logic should not be mixed directly.
How does GPT-4.1 maintain multi-turn conversations?
When using Chat Completions, put relevant history into messages; when using Responses, organize input and related conversation content according to the documentation. Provide the latest materials, revision goals, and key constraints each turn; for longer tasks, retain interim summaries and a final version that can be checked independently.
How can I make GPT-4.1 output processable JSON?
Chat Completions provides json_object and json_schema format settings. It is recommended to clearly specify field meanings, required fields, and rules for missing values, and perform structural and business validation on the receiving end. Valid JSON only means the format can be parsed; it does not mean amounts, classifications, or cited content are necessarily correct.