All models

tts-1-hd

OpenAIAudio
Get your API key
tts-1-hd

High-quality speech model for polished voice-overs and content narration

tts-1-hd is OpenAI's quality-first text-to-speech model, suitable for converting completed scripts, course copy, and product introductions into playable audio. Unlike tts-1, which focuses on real-time use, it is better suited to production workflows that prioritize the listening quality of the finished output. With preset voices, audio formats, and speech-rate options, text content can be integrated into dubbing, narration, and application announcements.

OpenAIModel brand
AudioModel type
AudioTask capability
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API host
api.acedata.cloud
model
tts-1-hd
Get your API key

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API features

Clarify capacity, inputs and outputs, and invocation methods before selecting a model.

Creation method
Text-to-speech, string input, binary audio output
Model positioning
Quality-first; tts-1 in the same series focuses on real-time use
Preset voices
alloy、echo、fable、onyx、nova、shimmer
Default voice
alloy
Audio formats
mp3、opus、aac、flac、wav、pcm; default: mp3
Speech-rate control
speed numeric parameter, default: 1.0
API endpoint
POST /v1/audio/speech; model is tts-1-hd

Quality-first is the native positioning of tts-1-hd; this platform provides access through POST /v1/audio/speech and supports the preset voices, audio formats, and speech-rate parameters listed above.

Core capabilities

Learn what tts-1-hd can bring to your work.

Choose based on the finished listening experience

The core trade-off of tts-1-hd is prioritizing speech quality rather than making real-time interaction its primary goal. For narration that has already been finalized and needs to be reviewed before release, use it after text proofreading and before audio editing, allowing the content team to choose based on the sound result without mixing speech generation and copywriting into the same step.

Preset voices support consistent production

The model provides six preset voices: alloy, echo, fable, onyx, nova, and shimmer. During production, first preview the same representative script, then assign one fixed voice to a course or content series. voice is used to select a preset timbre and is suitable for establishing a consistent narration configuration; it is not the same as uploading a real person's recording to replicate a voice.

Formats and speed support delivery

Audio can be selected as mp3, opus, aac, flac, wav, or pcm according to playback or production needs, with mp3 used by default; speed provides an option for adjusting speech rate. When web playback, mobile announcements, or post-production editing are needed, plan the file format and narration pace separately, and handle the result as audio binary data when receiving it.

Use cases

Start with specific tasks to find where the model can be effective.

Courses and knowledge explanations

Use proofread course scripts, concept explanations, or operating instructions as text input, and select a fixed voice to generate explanatory audio. Organize scripts by chapter, preview terminology, numbers, and transitions between paragraphs, then use approved clips in courseware or learning pages to create voice assets that can be reproduced as content is updated.

Product video narration

Input product introductions, feature demonstration commentary, or promotional video narration, compare preset voices with the visual style first, then choose a format for editing. The deliverable is speech audio for the input script; background music, sound effects, subtitles, and visual synchronization should be handled in the post-production workflow, rather than treating a voiceover request as full video production.

Article read-aloud and app announcements

Convert edited article paragraphs, help instructions, or app prompts into audio for readers to play on demand or within an interface. For fixed content that can be produced in advance, generate and review it first, then publish it with the page; for applications that need live responses, separate answer generation from answer narration into different tasks.

How to choose this model

Choose based on task complexity, input materials, and expected results.

Prioritize HD for finished voiceovers

If the task is to create courses, article readings, or video narration that can be played repeatedly, the quality-first positioning of tts-1-hd better matches the selection goal. When comparing it with tts-1, focus on the listening experience of representative scripts and the production workflow, rather than inferring sample rate, channels, or output duration solely from the HD name; preview actual content before formal production.

Choose separately for real-time interaction and transcription

If the core requirement is real-time voice feedback, the native positioning of tts-1 is closer to this direction, but the actual interactive experience still needs to be tested with the application. If the requirement is to turn existing recordings into text, choose an audio transcription model. tts-1-hd is responsible for reading text aloud; it does not replace speech recognition, generate conversational answers, or perform application operations.

Getting started: creating explanatory narration that needs repeated review

Prepare the input first, then connect it to the corresponding application workflow.

Prepare the input

Prepare proofread course or product narration, split the script according to visual segments, and choose a consistent voice.

Organize the call and follow-up workflow

Use /v1/audio/speech to submit model=tts-1-hd, input, and voice. Save the binary result according to the returned audio format, preview it in the actual player first, then incorporate the chosen voice and format into content production.

Practical task example: creating explanatory narration that needs repeated review

Design the task directly from the following inputs and acceptance priorities.

Recommended task

Select tts-1-hd through /v1/audio/speech; first create a representative audio sample containing terminology, numbers, and long sentences.

Key checks

Use headphones to assess clarity, noise, and pauses, then compare it with tts-1; even quality-first models require segment-by-segment review and editing.

Usage boundaries

Understand the output quality and scope of capabilities before formal use.

  • The basic input for tts-1-hd is text to be read aloud, and voice selection comes from a preset list. Do not treat voice as a real person's identity, a reference recording, or a voice-cloning instruction; if a project must reproduce a specific person's vocal timbre, choose another speech solution that explicitly supports that creation method.
  • Quality priority does not mean automatically completing professional voiceover review for every script. Proper nouns, abbreviations, numbers, and complex punctuation should be previewed before release; when adjustments are needed, first revise the text phrasing, segmentation, or speaking speed, then regenerate the corresponding segment.
  • The output is audio speech, not a complete program with subtitles, word-level timelines, and background music. Long content should be managed by production unit, with saving, playback, and stitching handled in the application; do not assume fixed audio duration or specific encoding parameters solely based on the HD name.

Frequently asked questions

Answers to common questions when using tts-1-hd.

How should I choose between tts-1-hd and tts-1?

tts-1-hd emphasizes quality, while tts-1 emphasizes real-time use. For course dubbing, article read-alouds, and narration that needs review before publication, try HD first; for real-time feedback, focus on evaluating tts-1. It is recommended to compare using the same actual script rather than interpreting the model difference as a fixed multiple of speed or audio quality.

How do I choose a voice for tts-1-hd?

Use voice to select alloy, echo, fable, onyx, nova, or shimmer; the default is alloy. Preview each one using a short script containing terms, numbers, and ordinary narration, then decide which voice suits the content. Once selected, keep the same configuration within the same series to help maintain a consistent audio production style.

Can tts-1-hd use my own recordings to customize a voice?

The creation method here is to enter text and select a preset voice, and it should not be used as a workflow for customizing a voice from a reference recording. voice contains a voice name, not a file address or person description. When voice cloning is needed, choose a product that explicitly supports reference audio, and handle authorization and production requirements separately.

Can the generated audio be used for playback and editing?

You can choose mp3, opus, aac, flac, wav, or pcm depending on the use case; the default output format is mp3. After calling, save it as binary audio or pass it to a player rather than parsing it as a text response. For editing, you must also arrange visual synchronization, background music, and final export in production software.

How do I submit a script and adjust the reading speed?

Submit input text to POST /v1/audio/speech, use tts-1-hd for model, and set voice, response_format, and speed as needed; the default value of speed is 1.0. In actual production, first preview the speaking speed with a short script, then process the main text; the text should also retain reasonable punctuation and paragraph boundaries.