English | 中文
A curated dataset hub for gastrointestinal (GI) endoscopy artificial intelligence, covering colonoscopy, wireless capsule endoscopy, laparoscopic/surgical endoscopy, polyp detection, segmentation, classification, video understanding, medical visual question answering, and multimodal clinical reasoning.
This repository is designed to help researchers quickly identify suitable datasets, understand their annotation types, locate official resources, and design fair and reproducible experiments.
Important: This repository does not host, redistribute, or re-license any original medical datasets. Always follow the license, data-use agreement, ethics approval, and institutional restrictions of the original dataset providers.
- Overview
- Collection Scope
- Datasets
- How to Use This Repository
- Repository Structure
- Citation
- Disclaimer
- Acknowledgements
GI endoscopy AI has evolved from early image classification, polyp detection, and polyp segmentation to video understanding, capsule endoscopy analysis, surgical scene perception, medical visual question answering, and multimodal clinical reasoning.
However, GI endoscopy datasets are scattered across papers, institutional websites, challenge platforms, GitHub, Kaggle, Zenodo, Hugging Face, OSF, Figshare, Synapse, PhysioNet, and other sources. They also differ greatly in task definition, modality, annotation granularity, access status, license, data split, and clinical applicability.
This project aims to answer questions such as:
- Which datasets are available for GI endoscopy AI?
- Which datasets are suitable for polyp detection, segmentation, classification, VQA, or multimodal learning?
- Which datasets provide masks, bounding boxes, point annotations, QA pairs, video labels, or clinical metadata?
- Which datasets are publicly available, require registration, require application, or may no longer be directly accessible?
- Which datasets are appropriate for baseline training, external validation, cross-dataset generalization, or medical multimodal model evaluation?
- How can researchers avoid data leakage, dataset overlap, and unfair evaluation?
This repository focuses on datasets related to GI endoscopy AI, including but not limited to:
This repository focuses on datasets related to artificial intelligence in gastrointestinal endoscopy, including but not limited to:
- Colonoscopy datasets: Image or video datasets for colonoscopy analysis.
- Polyp detection datasets: Datasets with bounding boxes, localization labels, or object detection annotations.
- Polyp segmentation datasets: Datasets with pixel-level, instance-level, or region-level masks.
- Polyp classification datasets: Datasets for benign/malignant classification, histological classification, or lesion-type recognition.
- Capsule endoscopy datasets: Wireless capsule endoscopy image or video datasets.
- Surgical endoscopy datasets: Datasets for laparoscopy, surgical tool recognition, surgical phase recognition, or surgical scene understanding.
- Endoscopy video datasets: Datasets for temporal modeling, object tracking, frame-level prediction, video-level classification, or dynamic lesion analysis.
- Medical VQA datasets: Image-text datasets built from GI endoscopy images and medical questions.
- Multimodal reasoning datasets: Datasets involving images, text, point annotations, localization, counting, reasoning chains, or clinical metadata.
- GI-related pathology datasets: Pathology or microscopy datasets closely related to GI lesions, polyps, tumors, or histological analysis.
The datasets are listed in reverse chronological order. Detailed dataset cards are placed under the corresponding dataset folders.
| Year | Dataset | Scale / Modality | Annotation | Resources |
|---|---|---|---|---|
| 2026 | InfoColon | 171K colonoscopy video frames; SNUH: 151 videos / 110K frames; CNUH: 13 videos / 9K image frames | Multi-class labels | Dataset | Code | DOI |
| 2026 | GPolypOD | 5,460 gastroenteroscopic polyp images | Bounding boxes / multi-label annotations | Dataset | DOI |
| 2025 | VIM-Polyp | 202 colonoscopy videos, 1,903 pathology images, IHC protein expression data, and clinical metadata | Multi-label annotations / clinical metadata | Dataset | Code | DOI |
| 2025 | Polyp-Size | 42 white-light colonoscopy videos | Polyp size annotations | Dataset | Code | DOI |
| 2025 | MedMultiPoints | 10,600 images, including GI endoscopy images and microscopy sperm images | Bounding boxes / points / counts / QA pairs | Dataset | Code | DOI | arXiv |
| 2025 | Kvasir-VQA-x1 | 6,500 original GI endoscopy images and 159K complex QA pairs | QA pairs | Dataset | Code | arXiv | DOI |
| 2025 | Galar | 80 capsule endoscopy videos, 350K annotated frames, and 29 labels | Multi-label annotations | Dataset | Code | DOI |
| 2025 | CAS-Colon | 78 colonoscopy withdrawal videos with 10 anatomical-region classes | Multi-class labels | Dataset | Code | DOI |
| 2024 | SEE-AI | 18,481 small-bowel capsule endoscopy images; 12 clinical abnormalities; 12,320 images annotated with 23,033 lesion boxes | Bounding boxes | Dataset | DOI |
| 2024 | REAL-Colon | 60 colonoscopy videos, 2.75M image frames, 132 resected colorectal polyps, and 350K frame-level bounding boxes | Bounding boxes | Dataset | Code | DOI | arXiv |
| 2024 | PolypDB | 3,934 static polyp images extracted from colonoscopy videos | Bounding boxes / masks | See dataset card |
| 2024 | Kvasir-VQA | 6,500 GI endoscopy images and 58,849 QA samples | QA pairs | Dataset | Code | arXiv | DOI |
| 2024 | HTPolypC | 202 colonoscopy videos, including 59 videos of hyperplastic polyps and tubular adenomatous polyps | Binary labels | Dataset | Code | DOI |
| 2024 | ERCPMP | 191 patients, 796 colonoscopy images, and 21 colonoscopy videos | Multi-label annotations | Dataset | arXiv | DOI |
| 2023 | PS-NBI2K | 2,000 NBI colonoscopy polyp images with pixel-level masks | Masks | Dataset | Code | DOI |
| 2023 | PolypGen | 8,037 colonoscopy frames; 3,762 positive frames and 4,275 negative frames | Bounding boxes / masks | Dataset | Code | arXiv | DOI |
| 2023 | MEDVQA-GI | 3,000 GI endoscopy images with QA pairs | QA pairs | Dataset | Code | DOI |
| 2023 | GastroVision | 8,000 GI endoscopy images with 27 image-level classes | Multi-label annotations | Dataset | Code | arXiv | DOI |
| 2023 | Endoscapes | 201 laparoscopic cholecystectomy videos; 11,090 frames with CVS assessment and 1,933 frames with bounding boxes | Bounding boxes / masks | Dataset | Code | arXiv | DOI |
| 2022 | CholecT50 | 50 laparoscopic cholecystectomy videos, about 100K frames, and 161K triplet instances | Multi-label annotations | Dataset | Code | arXiv | DOI |
| 2022 | SUN-SEG | 1,106 short video clips and 158K colonoscopy frames | Masks | Dataset | Code | arXiv | DOI |
| 2021 | SUN-Dataset | 152K colonoscopy frames; 49K polyp-positive frames and 103K polyp-negative frames | Bounding boxes | Dataset | Code | DOI |
| 2021 | LDPolypVideo | 160 fully annotated colonoscopy videos, 40K frames, 34K positive frames, and 200 annotated polyps | Bounding boxes | Dataset | Code | DOI |
| 2021 | Kvasir-Instrument | 590 GI endoscopy tool images | Bounding boxes / masks | Dataset | Code | arXiv | DOI |
| 2021 | KvasirCapsule-SEG | 55 VCE polyp images | Bounding boxes / masks | Dataset | Code | arXiv | DOI |
| 2021 | Kvasir-Capsule | 117 VCE videos, about 47K medically verified annotated images, and a large number of unlabeled video frames | Multi-label annotations / bounding boxes | Dataset | Code | DOI |
| 2021 | FCBUOR | 155 video sequences and 37K colonoscopy frames | Bounding boxes / binary labels | Dataset | arXiv | DOI |
| 2021 | BKAI-IGH-NeoPolyp-Small | 1,200 colonoscopy images, including 1,000 training images and 200 test images | Masks / multi-label annotations | Dataset | Code | arXiv | DOI |
| 2021 | PICCOLO | 3,433 colorectal polyp images, including 2,131 white-light images and 1,302 NBI images | Bounding boxes / masks / multi-label annotations | Dataset | DOI |
| 2020 | Kvasir-SEG | 1,000 colonoscopy polyp images with corresponding pixel-level masks | Masks | Dataset | Code | arXiv | DOI |
| 2020 | HyperKvasir | 110K GI endoscopy images, 374 videos, and 23 labeled image classes | Masks / multi-label annotations | Dataset | Code | DOI |
| 2020 | EDD2020 | 386 GI endoscopy images with 1,251 annotated objects from 10 classes | Bounding boxes / masks / multi-label annotations | Dataset | Code | arXiv | DOI |
| 2020 | CP-CHILD | 9,500 pediatric colonoscopy RGB images, including 1,400 polyp images and 8,100 non-polyp images | Binary labels | Dataset | DOI |
| 2020 | CholecTrack20 | 20 complete surgical videos, about 35K annotated frames, and 65K tool-instance labels | Bounding boxes / multi-label annotations / trajectories | Dataset | Code | arXiv | Paper |
| 2017 | Nerthus | 21 GI endoscopy videos with 5,525 frames | Multi-class labels | Dataset | DOI |
| 2017 | Kvasir | 4,000 GI endoscopy images from 8 classes, with 500 images per class | Multi-class labels | Dataset | DOI |
| 2017 | CVC-EndoSceneStill | 912 colonoscopy images from 44 videos of 36 patients | Masks | Dataset | DOI |
| 2016 | Colonoscopic | 76 lesion samples, including 40 adenomas, 21 hyperplastic lesions, and 15 serrated adenomas | Multi-class labels | Dataset | DOI |
| 2015 | CVC-ClinicDB | 25 clinical videos and 29 polyp-containing GI endoscopy videos, with 612 still images | Masks | Dataset | DOI |
| 2012 | CVC-ColonDB | 300 polyp-containing images from 15 colonoscopy videos, covering different views, distances, and morphologies | Masks | Dataset | DOI |
A recommended workflow is:
- Select the task. Start from detection, segmentation, classification, VQA, video understanding, surgical scene analysis, or multimodal learning.
- Check annotation granularity. Confirm whether the dataset provides image-level labels, masks, bounding boxes, points, QA pairs, video labels, or clinical metadata.
- Verify the official source. Follow the dataset provider's latest download instructions and license terms.
- Inspect data splits. Use official splits when available. Avoid mixing frames from the same video or patient across training and testing.
- Use external validation. For robustness studies, evaluate models on independent datasets or cross-center datasets whenever possible.
- Report dataset details. Clearly report version, access date, split strategy, preprocessing, and evaluation metrics.
Awesome-GI-Endoscopy-Datasets/
├── README.md # English homepage
├── README_zh.md # Chinese homepage
├── datasets/ # Dataset cards and dataset-specific guides
│ ├── 2026-InfoColon/
│ ├── 2025-Kvasir-VQA-x1/
│ ├── 2024-REAL-Colon/
│ └── ...
├── docs/ # Methodological notes, task guides, and benchmark suggestions
├── assets/ # Figures, logos, and visual examples
├── scripts/ # Utility scripts for parsing, checking, and visualization
└── configs/ # Path templates and reusable configuration files
If this repository is useful for your research, teaching, experiment design, or dataset selection, please cite or acknowledge it.
@misc{awesome_gi_endoscopy_datasets,
title = {Awesome GI Endoscopy Datasets: A Curated Dataset Hub for Gastrointestinal Endoscopy AI},
author = {CloneIQ and contributors},
year = {2026},
howpublished = {\url{https://github.com/cloneiq/Awesome-GI-Endoscopy-Datasets}},
note = {GitHub repository}
}This repository does not host, redistribute, or re-license any original medical imaging dataset.
It only provides:
- dataset metadata;
- dataset cards;
- official links;
- paper links;
- download notes;
- license notes;
- usage suggestions;
- benchmark and evaluation suggestions.
All datasets remain governed by their original licenses, data-use agreements, ethics approvals, and institutional restrictions. Users must verify and comply with the latest terms from the original dataset providers before downloading, using, modifying, or redistributing any dataset.
The information in this repository is provided for research and educational purposes only. Dataset availability, licenses, links, and documentation may change over time. Always refer to the latest official source.
We thank all researchers, clinicians, medical institutions, challenge organizers, dataset maintainers, and open-source contributors who have built, released, maintained, and evaluated GI endoscopy datasets.
This repository will continue to serve as an open community resource for reliable, reproducible, interpretable, and clinically meaningful GI endoscopy AI research.
Lijun Liu, Associate Professor (Ph.D.), Kunming University of Science and Technology Kunming, Yunnan CHINA, email: cloneiq@kust.edu.cn