Which invocation ID should Gemini 3.7 Flash use?
Use Chat Completions, set model to gemini-3.7-flash, organize text and images in messages, and read responses from choices; for streaming interactions, handle increments according to the documentation. The capabilities and message format of native Generate Content must be verified independently; do not infer that the platform exposes all features from the vendor's native specifications.
How much input and output does it support?
The native input limit is 1,048,576 tokens, and the output limit is 65,536 tokens; they are not the same budget. When organizing long tasks, you should still select relevant materials and set response length based on the deliverable; these native numbers do not mean the full capacity should be used for every call.
How do I set its thinking intensity?
You can choose low, medium, or high; minimal is not supported natively. Simple summaries can start with low, while multi-condition analysis can try medium or high. Do not treat an extremely small response budget as a way to disable thinking; the Gemini 3.x Flash guide recommends setting max_tokens above 512.
Can it read PDFs and generate images or audio?
Prepare document text, table data, or clear page screenshots relevant to the question, and specify whether you need a summary, comparison, or extraction of particular information. Submit content in formats supported by the selected public interface; a PDF address cannot be used as image_url. Require results to retain original-text locations, field sources, and unconfirmed items, and verify key numbers against the source materials.
How do I continue a conversation after calling tools?
The application maintains the Chat Completions messages history, including relevant user and assistant messages, the latest materials, and feedback. For longer tasks, periodically consolidate interim conclusions and conditions that must be retained before continuing to ask questions; continuous conversation does not mean unlimited memory.