SNLI Dataset

SNLI Natural Language Inference
Dataset

The first large-scale natural language inference dataset, with 570,152 sentence pairs containing entailment, contradiction, neutral, and a small number of undecided labels; a foundational benchmark in the NLI field.

570,152 sentence pairs 3 label types CC BY-SA 4.0 License Bowman et al. (2015)
SNLI Dataset
📊
570,152
Total Sentence Pairs
🏷️
3
Label Categories
👥
1,119
Undecided Majority Labels
📜
CC BY-SA 4.0
Open License Agreement

Dataset Highlights

A foundational benchmark dataset in the NLI field, setting the standard for advancing natural language understanding research

🌍

Large-Scale Human Annotation

The official 1.0 release contains 550,152 training, 10,000 development, and 10,000 test examples; some examples have only one annotator, and 1,119 examples did not reach a majority label.

🔗

Three Types of Semantic Relations

Covers the three core inference relations of entailment, contradiction, and neutral, spanning the fundamental dimensions of natural language understanding.

🎯

Foundational Benchmark

As the first large-scale NLI dataset, SNLI has been widely used to evaluate and compare various natural language understanding models, driving the development of models such as BERT and GPT.

📷

Visually Grounded Text

Most premise sentences come from Flickr30k image captions, while approximately 4,000 training sentences come from VisualGenome pilot data; the distribution package includes source documentation.

📖

Annotator Agreement

The package retains official per-example annotation fields; the number of annotators varies across examples, so not all records can be treated as having five-annotator consensus.

🏛️

Authoritative Academic Source

Released by Stanford University's NLP Group, the paper has been cited more than ten thousand times and is one of the most influential datasets in the NLP field, widely adopted in academia and industry.

Use Cases

From foundational research to industrial applications, covering the full natural language understanding pipeline

🧠

Natural Language Inference

Train and evaluate NLI models to determine entailment, contradiction, or neutral relationships between two sentences

📐

Sentence Embeddings

Use sentence-pair relationships to train high-quality sentence vector representations, improving semantic similarity and retrieval performance

🔄

Transfer Learning

Fine-tune models such as BERT and RoBERTa using it as a pretraining task to improve downstream NLU task performance

✅

Textual Entailment Detection

Build core inference modules for applications such as fact verification, question-answering systems, and text consistency checking

Natural Language Inference Textual Entailment Sentence Understanding Stanford NLP Benchmark Dataset

Data Preview

The following are official SNLI 1.0 source file statistics; unreviewed original sentences are not disclosed.

Summary
SNLI 1.0 sentence pairs: 570,152
Training set: 550,152
Development set: 10,000
Test set: 10,000
Undetermined majority labels (-): 1,119
Delivery format: ZIP containing original JSONL and license and attribution files

Get Started Quickly in 3 Steps

From browsing to research, start your NLI experiments in minutes

01

Browse the Dataset

View dataset details on the Ace Data Cloud platform to learn about metadata such as field descriptions, label distributions, and license agreements.

02

Download Data After Purchase

Download a private ZIP after purchase: includes official train / dev / test JSONL files (570,152 entries total), README, Flickr30k / VisualGenome source information, and CC BY-SA 4.0 license text.

03

Load and Train

After extracting the ZIP, read the three JSONL files; datasets.load_dataset("snli") separately loads from a public repository and does not read the purchased delivery package.

Start Exploring Natural Language Inference Data

A foundational benchmark in the NLI field. Review the real summary before purchasing, then download a delivery package with source licenses after purchase. Whether you are an NLP researcher or a deep learning engineer, SNLI is an indispensable foundation for experimentation.