Convert recordings into original-language text and usable subtitles
whisper-1 is OpenAI's speech-to-text model, corresponding to Whisper large-v2 at the time of the official API release. It is suitable for organizing speech from interviews, courses, voice messages, and audio-visual materials into original-language text, with text, JSON, or subtitle formats selected according to the workflow. On this platform, you can submit files through the audio transcription entry point to integrate recording processing into content production and document organization workflows.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Clarify capacity, inputs and outputs, and calling methods before selecting a model.
Native model relationship
Whisper large-v2; API call ID is whisper-1
Input and output
Audio file input, original-language text transcription output
Native file formats
m4a、mp3、mp4、mpeg、mpga、wav、webm
API endpoint
POST /v1/audio/transcriptions; submit binary file
Output format options
json、text、srt、verbose_json、vtt; default is json
Transcription controls
language and prompt can be passed; temperature defaults to 0
Native format details and call parameters are listed separately; this entry point focuses on file transcription and is not equivalent to real-time voice conversation.
Core capabilities
Learn what whisper-1 can bring to your work.
A transcription path that preserves the original language
whisper-1's transcription task converts spoken audio into original-language text rather than translating it into English by default. This is suitable for interviews, courses, and dictated materials where original wording needs to be preserved. The generated text can serve as a basis for subsequent search, excerpting, and editing, while summarization and rewriting should be handled as separate processing steps.
Deliver text and subtitles separately
The same transcription workflow can select output formats based on use: text is suitable for direct entry into an editor, JSON makes text easier for programs to read, and SRT and VTT are suitable for subtitle production. Format selection should be designed around the final deliverable to avoid first saving plain text and then separately rebuilding subtitle structure and processing workflows.
Organize calls around files
Calls center on audio files, so recordings do not need to be converted into chat messages. Requests can include language and prompt text to help applications organize transcription context; prompts should be prepared around the recording content rather than treated as free-form creative instructions. This makes it easier to connect uploading, transcription, saving, and manual proofreading into a stable workflow.
Use Cases
Start with specific tasks to identify where the model can be useful.
Interview and Oral History Organization
Submit interview recordings or oral materials as files, choose text or JSON output, and receive transcripts that are easy to search and edit. Editors can use them to locate topics, organize quotations, and listen back to verify names, organization names, and key original statements; article structure and insight extraction can be handled separately after transcription is complete.
Course and Video Subtitle Production
Choose SRT or VTT output for course recordings or video materials, then send the results into a subtitle editing or playback workflow. Before delivery, check sentence breaks and the correspondence between text and timing against the original material; the model handles speech transcription, while subtitle layout, visual styling, and final video export are still completed by production tools.
Voice Message Archiving
Convert user voice messages into text for storage in ticketing systems or knowledge bases. Applications can read the text in JSON and associate the transcription with the original audio, making it easier for staff to review and listen back. Classification, priority assessment, and response generation should be handled in separate processing steps rather than treating transcription results directly as business decisions.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
Choose It When You Need a Text Transcript
If the core requirement is to process existing recordings and deliver original-language text or subtitles, whisper-1 has a clear task boundary. It is neither a general question-answering model nor a speech synthesis model: it will not read text aloud as audio, nor should it be expected to handle all of the work of automatically writing meeting notes. Complete transcription first, then follow with summarization or editing workflows for easier verification against the original speech.
Differentiate by Version and Task
whisper-1 is the API call name, while large-v2 is the corresponding native version from the official release; it should not be understood as a blanket term for all newer Whisper versions. When comparing other transcription models, use the same recording to check proper nouns, sentence breaks, and required output formats; if the goal is English translation, also distinguish between transcription in the original language and a separate translation task.
Get Started: Create Editable Subtitles for an Interview
Arrange the input first, then connect it to the appropriate application workflow.
Prepare the Input
Prepare the interview audio and correct proper nouns, and specify whether the output should be an original-language transcription or subtitles.
Organize the Call and Follow-up Workflow
Upload the binary file through /v1/audio/transcriptions and explicitly select whisper-1; first obtain the basic text result, then choose verified subtitle or timing information output as needed for the task. Keep the original recording available for playback during editing.
Practical Task Example: Creating Editable Subtitles for Interviews
Design the task directly from the following inputs and acceptance priorities.
Suggested Task
Upload audio using file, with model=whisper-1; when subtitles are needed, select the corresponding srt or vtt format in the documentation, and include language and terminology prompts.
Key Checks
Verify text, sentence breaks, and timing alignment in the player, and listen again to overlapping speech from multiple speakers; file transcription cannot be directly equated with real-time voice conversations.
Usage Boundaries
Before formal use, understand the output quality and scope of capabilities.
Transcribed text is not automatically generated summaries, meeting conclusions, or speaker identity records. When multi-person recordings require speaker differentiation, arrange a separate speaker-labeling step; do not infer the attribution of each sentence solely from the transcription text, and maintain the correspondence between important quotes and the original audio.
Use audio files as input; do not submit PDFs, chat text, or web links directly as binary file. When calling through MCP's audio URL tool, distinguish this from the HTTP file upload flow as well; file preparation and reading methods depend on the client used.
File transcription is not the same as a real-time speech service that recognizes speech while recording, and subtitle output does not mean that all editing work is complete. When real-time interaction, precise timing alignment, or formal subtitle delivery is required, test the workflow using actual audio and schedule review listening and subtitle editing.
Frequently Asked Questions
Answers to common questions about using whisper-1.
What is the relationship between whisper-1 and Whisper large-v2?
whisper-1 is the model name used when calling the API, and OpenAI's Whisper API release notes map it to large-v2. They describe the invocation name and the native version respectively. whisper-1 should not be treated as a generic alias for all Whisper versions, nor should this be taken to mean it is the latest version.
Will Chinese recordings automatically become English?
The audio transcription endpoint outputs text in the original language; it does not translate to English by default. The official documentation distinguishes transcription in the original language from translation into English as separate tasks. Therefore, Chinese recordings should be processed using the Chinese transcription workflow; if an English transcript is needed, a separate translation step should be arranged.
Can SRT or VTT subtitles be generated directly?
The API parameters provide srt and vtt output options for subtitle workflows; for plain-text tasks, text or json can be selected. After subtitles are generated, the content, line breaks, and timing should still be checked against the audio before handing them to subtitle editing or playback tools; they should not be considered a finished production directly.
Can prompt make it rewrite a recording into an article?
prompt is hint text in a transcription request and should not be treated as a general-purpose writing instruction. The core deliverable of whisper-1 remains an audio transcript. If an article, summary, or action items are needed, it is recommended to preserve the original transcription result and then pass it to a separate text-processing step, avoiding confusion between the original words and processed content.
Should I upload a file or submit an audio URL?
When calling /v1/audio/transcriptions directly, you need to submit a binary file and use whisper-1. If using an MCP client, you can choose its tool for transcribing audio from a URL. The input organization differs between the two approaches; do not pass a URL string directly as HTTP's binary file field.