Efficient Text Embedding Model for Semantic Retrieval and Knowledge Bases
text-embedding-3-small is OpenAI's third-generation text embedding model, designed to convert text into numerical vectors that represent semantics rather than generate chat responses. It is suitable for knowledge base retrieval, similar content matching, and text clustering, and supports adjusting vector dimensions to balance retrieval performance, storage space, and computational overhead, making it a practical starting point for building everyday semantic retrieval systems.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and Interface Features
Clarify capacity, inputs and outputs, and invocation methods before choosing a model.
Model type
Text embeddings; outputs semantic vectors and does not generate responses
API endpoint
POST /v1/embeddings;model=text-embedding-3-small
Input formats
Non-empty text, arrays of text, arrays of non-negative integer tokens, or batches thereof
Batch size
Text batches or token array batches support up to 2048 items
Dimension control
dimensions can shorten vectors; full dimensions are output when unspecified
Output encoding
float or base64, with float as the default; results include index and token usage
Dimension shortening is a native model capability; the input formats, batch size, and output encodings above correspond to this platform's API endpoint.
Core Capabilities
Learn what text-embedding-3-small can bring to your work.
Connect different expressions with semantics
The model maps text content to numerical vectors, providing a foundation for similarity retrieval. Even when user questions and source materials use different wording, related content can be found through vector comparison. It is responsible for semantic representation; retrieval ranking, filtering criteria, and final answers are still handled by the application, making it suitable for use alongside keyword search.
Support multilingual retrieval tasks
In official benchmark results, text-embedding-3-small achieved an average MIRACL score of 44.0%, higher than ada-002's 31.4%; its average MTEB score was 62.3%. These results reflect retrieval improvements over the previous version, but specific languages, industry terminology, and text chunking methods can still affect business performance.
Adjust dimensions according to retrieval needs
Using dimensions can shorten output vectors, reducing the number of values per record and helping balance vector database storage and computational overhead. Full dimensions are output when it is not specified. After shortening, retrieval performance should be reevaluated, and document vectors and query vectors should use the same model and the same dimension configuration.
Use Cases
Start with specific tasks to find where the model can be effective.
Knowledge base retrieval
First extract product manuals, help documentation, or internal policies into text and split them into segments, then generate vectors in batches and write them to a retrieval database. When a question is asked, generate a vector for it, find relevant passages, and provide them to an answer model as reference. This model handles the document retrieval stage and does not directly produce knowledge base answers with citations.
Similar content search
Convert customer service tickets, issue titles, or product descriptions into vectors, compare the similarity of new content with existing records, and return a candidate list. This can be used to find similar incidents, link historical resolution cases, or identify similar descriptions; whether content is duplicate or can be merged should be determined using business rules and human review.
Text clustering and topic organization
Vectorize feedback, comments, or short texts in batches, then use clustering algorithms to organize content groups and deliver topic clusters and representative texts. Applications can use the returned index to align with original records, making it easier to review the content in each group. The model provides vector representations; topic naming and summarization must be completed separately.
How to choose this model
Choose based on task complexity, input materials, and expected results.
Evaluate small first for new systems
If the goal is everyday semantic search, similar-text matching, or knowledge-base retrieval, you can first use small to establish a baseline. Compared with ada-002, it performed better in the multilingual retrieval and English task evaluations at the time of its official release, and it supports dimension reduction. When migrating existing legacy indexes, regenerate document vectors instead of directly mixing results from different models.
Compare large when accuracy is the priority
If your business is more sensitive to retrieval accuracy, compare small and text-embedding-3-large on the same test set. The large model scored higher in the official release evaluations, but that does not mean every business will see the same benefit. It is recommended to combine real queries, relevant-passage annotations, and index configuration to determine whether the improvement is worth using a larger model.
Getting started: Build lightweight semantic search for an FAQ
Prepare the inputs first, then connect them to the corresponding application workflow.
Prepare inputs
Prepare a batch of FAQs, common questions, and synonymous rewrites, and define the knowledge update method and access scope.
Organize calls and subsequent workflows
Use /v1/embeddings to select text-embedding-3-small, and manage text and its document IDs separately. Write vectors to an index with fixed configuration, and use the same encoding for queries; first verify that candidate source text is traceable, then use the retrieval results in the application.
Practical task example: Build lightweight semantic search for an FAQ
Design the task directly from the inputs and acceptance priorities below.
Recommended task
Encode FAQs and queries with text-embedding-3-small, create a fixed-dimension index, and retrieve candidate entries using cosine similarity.
Key checks
Check whether different phrasings can find the same answer, and evaluate storage and query time; rebuild the corresponding index after changing the model or dimensions.
Usage boundaries
Before formal use, understand the output quality and scope of capabilities.
It outputs numerical vectors and does not directly generate natural-language responses, automatically build a vector database, run clustering, or perform retrieval. A complete application still requires steps such as text processing, storage, and similarity calculation; RAG scenarios also require a separate response model.
Input must be non-empty text or a compliant token array; PDF files cannot be submitted directly as text. A maximum batch size of 2048 items refers to the number of input items and does not mean documents of arbitrary length can be processed at once; for long documents, extract the main text first and split it semantically.
Smaller dimensions are not always better, nor should they be understood as allowing vectors to be expanded arbitrarily. After adjusting dimensions or changing models, update index configuration and document vectors accordingly; retrieval performance must be validated with real queries, and average benchmark scores cannot replace business acceptance testing.
Frequently asked questions
Answers to common questions about using text-embedding-3-small.
Can text-embedding-3-small answer questions directly?
No. It converts questions or materials into vectors for applications to compare semantic similarity. Knowledge-base Q&A typically uses it first to retrieve relevant passages, then passes them to a response model to compose an answer; calling only the embedding API returns vectors, not explanations, summaries, or conversational replies.
How does it differ from text-embedding-ada-002?
It is a third-generation embedding model, not an alias for ada-002. Official release benchmarks show that it achieves higher average scores on multilingual retrieval and English tasks, and it supports shortening vectors through dimensions. When migrating, rebuild document vectors to avoid comparing results from the two models in the same space.
When should text-embedding-3-large be chosen?
When retrieval precision matters more than lightweight operation, it is worth comparing large. It scored higher in the official benchmarks released at the same time, but the choice should still be based on your own query set and document repository. You can first establish a small baseline, then evaluate whether large improves retrieval of relevant passages for key questions.
How do I submit batches and match returned results?
Submit model and input to POST /v1/embeddings. input can use batches of text arrays or token arrays, with a maximum of 2048 items. Each item in the returned data contains index and embedding; use index to align with the original input, and check usage for token consumption.
What do dimensions and encoding_format control respectively?
dimensions controls vector shortening; if unspecified, the full dimensions are output. encoding_format controls the returned representation, with float or base64 available and float as the default. The former requires evaluating retrieval performance, while the latter is for adapting to data-processing methods; encoding choices should not be understood as model accuracy levels.