All models

text-embedding-ada-002

OpenAIEmbedding
Get your API key
text-embedding-ada-002

Classic text embedding model for existing knowledge bases and semantic retrieval

text-embedding-ada-002 is OpenAI's classic text embedding model, converting natural language and code into comparable numerical vectors for semantic retrieval, similar content matching, and code search. It connects queries and documents through a unified representation, making it suitable for maintaining existing ada-002 vector databases and serving as a baseline for retrieval experiments; the output is vectors, not chat responses.

OpenAIModel brand
VectorsModel type
VectorsTask capability
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API host
api.acedata.cloud
model
text-embedding-ada-002
Get your API key

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API features

Clarify capacity, input/output, and invocation methods before choosing a model.

Native input capacity
8192 tokens context length
Full vector dimensions
1536 dimensions
Input methods
Non-empty text, batches of text, token arrays, or batches of token arrays
Batch size
Up to 2048 items for text batches and token batches
Output encoding
float or base64, float by default
Invocation method
POST /v1/embeddings;model=text-embedding-ada-002

8192 tokens and 1536 dimensions are publicly available native specifications. Batch formats and return encoding are invocation settings for this platform entry point, and batch size is not equivalent to total token capacity.

Core capabilities

Learn what text-embedding-ada-002 can bring to your work.

Connect queries and documents with semantics

The model converts queries and documents into the same vector representation, making it suitable for finding content with different wording but similar meaning. Compared with retrieval methods that rely only on keywords, it can handle semantic candidate retrieval; applications can then combine keywords, business filters, or ranking rules to form final search results, rather than treating similarity directly as a factual judgment.

Unified text and code search

ada-002 unifies text similarity, query and document search, and natural-language and code search in one model. It can both index documentation and represent code snippets, supporting the search for relevant implementations by functional description. Its code capability here is semantic representation and retrieval, not writing programs, running code, or verifying program correctness.

Structured vectors for easy application integration

The complete output is a 1536-dimensional vector, making it convenient to configure indexes with fixed dimensions. Returned data includes an index, which can be used to associate inputs within a batch, and provides token usage information. The default float format is convenient for numerical processing; when base64 is selected, it must be decoded according to its encoding before being passed to vector computation or storage components.

Use cases

Start with specific tasks to find where the model can be effective.

Maintain retrieval for an existing knowledge base

Split new help documents, product descriptions, and internal policies into text chunks, continue using ada-002 to generate vectors, and write them to the existing index. User questions are encoded with the same model, and relevant passages are then retrieved and passed to a response model. The deliverable is traceable candidate documents and passages, not answers generated directly by the embedding API.

Find similar content and duplicate topics

Input ticket bodies, article summaries, or product descriptions, then rank or cluster the resulting vectors by similarity to find similar issues, consolidate duplicate topics, and recommend related content. Applications must set their own thresholds and sample-check results, especially distinguishing between similar topics and complete business duplicates, to avoid automatically merging important differences.

Build code snippet search

Encode function snippets, comments, and development documentation separately, then convert natural-language needs such as “parse configuration files” into query vectors to return relevant code locations and descriptions. This is suitable for codebase navigation and implementation reference; indexes should retain file paths and version information so developers can return to the original files and inspect the actual logic.

How to Choose This Model

Choose based on task complexity, input materials, and expected results.

Prioritize consistency when you already have an ada index

If your existing document vectors were generated by ada-002, continue using it for new content and queries to help keep the retrieval space consistent. Do not mix it directly into the old index just because another model also outputs the same dimensionality. When planning an upgrade, re-encode the documents and compare recall results, storage overhead, and migration costs using real queries before deciding whether to switch.

Evaluate third-generation embeddings for new projects as well

When building a new retrieval system, it is recommended to evaluate both text-embedding-3-small and text-embedding-3-large. Official comparisons show that both outperform ada-002 on multilingual retrieval and average benchmarks for English tasks, and they offer native vector shortening capabilities. ada-002 is better suited as a continuation of an existing system or an experimental baseline; the choice should still be based on testing with business data.

Getting Started: Safely Maintain an Existing ada Vector Database

First organize the inputs, then connect them to the corresponding application workflow.

Prepare inputs

Confirm the model used by the historical index, the 1536-dimensional configuration, text chunking, and the similarity algorithm.

Organize calls and subsequent workflows

Use /v1/embeddings to select text-embedding-ada-002, and manage text and its document IDs separately. Write vectors to an index with a fixed configuration, and use the same encoding for queries; first confirm that the original text behind candidates is traceable, then use the retrieval results in the application.

Practical Task Example: Safely Maintain an Existing ada Vector Database

Design tasks directly around the following inputs and acceptance priorities.

Recommended task

Continue using text-embedding-ada-002 to encode new documents and queries, maintaining the same processing workflow; create a separate index for evaluating new models.

Key checks

Verify vector dimensionality, index compatibility, and historical retrieval examples; do not mix vectors from different models into the same space or merely truncate old vectors instead of re-encoding them.

Usage Boundaries

Before formal use, understand the output quality and capability scope.

  • 8192 tokens is the input capacity, not the number of Chinese characters; long texts should be counted first and split at content boundaries. A batch maximum of 2048 items also does not mean that every item can simultaneously use the full native capacity. Reasonable chunking not only helps control request size, but also allows retrieval results to locate more specific passages.
  • Do not use ada-002 as a model that supports native variable dimensions; conventional indexes should be designed for the full 1536 dimensions. Arbitrarily truncating vectors cannot be considered an equivalent substitute; if the target vector database requires fewer dimensions, consider the text-embedding-3 series, which has native shortening capabilities.
  • Retrieval capability does not mean better performance on all classification tasks. Official sources state that it did not outperform text-similarity-davinci-001 on the SentEval linear-probe classification benchmark; if training a lightweight classification layer on vectors, validate label discrimination separately rather than applying conclusions from search tasks.

Frequently Asked Questions

Answers to common questions about using text-embedding-ada-002.

Can ada-002 directly answer knowledge base questions?

It cannot generate answers directly. It converts questions and documents into vectors, which the application uses to find relevant passages, and then a generative model organizes the answer. When building knowledge base question answering, document storage, similarity retrieval, and answer generation must be completed separately; the embeddings API handles the text representation part.

How can I generate vectors in batches and map them to the original text?

Submit an array of non-empty strings to input, or submit a batch made up of token arrays; both batch formats support up to 2048 items. Each item in the response data includes an index, which can be used to associate it with the input. It is recommended to also save business document IDs so that the original content can still be found after vectors are written to the index.

Can I directly submit PDFs, images, or web addresses?

This endpoint accepts text or token arrays; it is not a file parsing API. Text must first be extracted from PDFs, image content must first be converted into searchable text, and the main content must first be retrieved from web pages. Even if a URL is submitted as a string, it does not mean the model will automatically open the page and read its content.

Do float and base64 change vector dimensions?

They are used to select the return encoding, not the model or dimensions. float returns an array of numeric values, suitable for direct vector processing; base64 returns an encoded string, which must be decoded before use. The full vector for ada-002 has 1536 dimensions; a change in transmission format should not be understood as vector shortening.

Do I need to rebuild the index when switching to text-embedding-3-small?

You should re-encode the documents and build a corresponding index, and queries must also use the same new model. Vectors produced by different models should not be mixed for comparison merely because they have the same dimensionality. Before migration, you can keep the old index for comparison, use real search queries to check result quality, and then gradually switch the retrieval workflow.