Add expressiveness to dialogue, narration, and brand voices
Fish Audio S2 Pro is a text-to-speech model for natural speech synthesis and voice cloning, suitable for commercial narration, character dialogue, and content that requires a distinct expressive rhythm. Platform guidance emphasizes its expressiveness and sets it as the default model for this entry point. By unifying text, reference voices, and prosody settings, developers can incorporate voiceover production into applications and content workflows.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Clarify the model, inputs and outputs, and integration method before selecting.
mp3, wav, pcm; both wav and pcm return a WAV container
Voice reference
reference_id or a single references sample; the two are mutually exclusive
Prosody control
prosody.speed controls speech rate, and prosody.volume controls volume gain
The S2 series can use [bracket] natural-language expression prompts in text, such as [whisper]; actual results should be confirmed by listening. Capability descriptions are based on public model materials and this platform's documentation; actual parameters, outputs, and billing rules are subject to the corresponding API and pricing sections.
Core capabilities
Learn about the voice tasks that s2-pro is suited to solve.
Focus on the expressive rhythm of speech
Punctuation, pauses, and sentence structure affect vocal performance. S2 Pro is suitable for comparing narration styles in short samples: first confirm emphasis and rhythm, then expand to the full script; the model's expressive positioning does not mean it can precisely execute every emotional marker.
Build a sense of character from reference voices
You can either reuse a saved voice through reference_id or use a recording and accurate verbatim transcript for one-time cloning. The former is suitable for recurring characters, while the latter is suitable for one-off voiceover projects; choose one of the two methods.
Balance auditioning and application delivery
Use MP3 for web playback or quick review, and use a WAV container for editing. After synchronous completion, retrieve the result through audio_url; longer tasks can also be configured with callbacks to integrate the waiting process into a task list.
Use Cases
Choose based on specific content and delivery method.
Short Videos and Advertising Voiceovers
Break product selling points, transitions, and endings into short segments, and compare how speaking speed and script wording affect delivery. Produce the full clip after confirming the result, then check duration and pauses against the visuals.
Character Lines and Interactive Content
Choose a voice that can be repeatedly used for a fixed character, and check consistency around actual lines. The application handles character arrangement and audio stitching; a single TTS request cannot directly be equated with automatic multi-character production.
Brand Narratives and Product Demonstrations
Synthesize feature descriptions, brand stories, and guiding copy into a unified voice. Preserve the correspondence between preview scripts and final clips, making it easy to replace sections after copy changes without remaking the entire video.
How to Choose This Model
Compare based on scripts, voice quality, and production cost.
Default Models Should Also Be Clearly Recorded
When the model request header is not set, the platform uses s2-pro. For production, it is still recommended to explicitly specify the model to avoid mixing different models between test scripts and subsequent content due to setting changes.
Compare Expressiveness and Next-Generation Capabilities
You can compare S2 Pro and S2.1 Pro using the same line and voice, focusing on speech naturalness, generation wait time, and the amount of subsequent revision. When long-form stability matters more, add S1 as a candidate.
Getting Started: Create a Character Line with Expression Prompts
Arrange the inputs first, then connect them to the corresponding application workflow.
Prepare Inputs
Prepare a short line, a character voice, and two emotional directions; first select an approved reference_id.
Organize the Request and Follow-Up Workflow
Submit text to /fish/tts, select s2-pro through the model request header, and provide either a reference_id or a single recording reference. After completion, preview with audio_url; for serialized content, save the model, voice, speaking speed, and format.
Practical Task Example: Create a Character Line with Expression Prompts
Design the task directly from the inputs and acceptance priorities below.
Suggested Task
Specify model: s2-pro through /fish/tts; you can compare the effects of natural-language expression prompts such as [whisper] in text with ordinary scripts.
Key Checks
Preview whether the delivery sounds natural, whether markers are read aloud, and whether pronunciation is complete; the single voice reference in this entry point does not equal automatic multi-character arrangement.
Usage Boundaries
Understand the synthesis method and delivery scope.
This endpoint is for text-to-speech and is not equivalent to speech recognition, music generation, or native real-time audio streaming. Long-form content should be produced in segments and reviewed, checking numbers, abbreviations, and proper nouns.
One-time cloning accepts only one publicly accessible HTTPS MP3/WAV recording sample and its accurate verbatim transcript; Base64, data URIs, or URLs with credentials are not accepted. It does not automatically save a long-term voice.
format=pcm returns a WAV container and cannot be processed directly as raw PCM bytes; opus is not supported. For the supported range of generation parameters and actual billing, see this platform's API documentation and pricing.
Frequently Asked Questions
Answers to common questions when using s2-pro.
Is S2 Pro the platform's default model?
Yes. This platform's Fish TTS endpoint defaults to s2-pro when the model request header is not set. We recommend explicitly recording the model for production tasks; the latest model officially recommended and this endpoint's default model are two different pieces of information.
Where should the model be specified?
Specify s2-pro through the HTTP request header model. It defaults to s2-pro when not specified; the voice reference_id is for voice selection and is configured separately from the synthesis model.
How should I choose between saving a voice and one-time cloning?
For fixed characters or series content, you can reuse a voice with reference_id; for one-time projects, you can use references to provide a recording and transcript. The two are mutually exclusive, and one-time cloning accepts only one reference sample.
How do I control speech rate and output format?
prosody.speed=1.0 indicates the original speed, volume uses dB, and 0 means the volume is unchanged. You can choose mp3, wav, or pcm; the latter two both use WAV containers, and MP3 bitrates can be 64, 128, or 192.
How do I track completion results for long scripts?
After setting callback_url, first save task_id and started_at, wait for the completion callback, or query by task ID. Only after obtaining audio_url can you proceed to playback or editing.