HEST-1k: A Dataset for Spatial Transcriptomics and Histology Image Analysis
In June 2024, the HEST-1k team released a dataset containing 1,229 spatial transcriptomics profiles, each associated with H&E whole-slide images and metadata. The data comes from 153 cohorts, covering 26 organs, 2 species, and 25 cancer types, including 2.1 million expression-morphology pairs and 76 million nuclei. The accompanying HEST-Library and HEST-Benchmark are used for foundation model evaluation.
HEST-1k is currently the largest paired dataset of spatial transcriptomics and histology images, addressing the scarcity of data and lack of standards in this field. It supports benchmarking of pathology foundation models, biomarker exploration, and multimodal representation learning, potentially accelerating precision medicine and drug development.
HEST-1k integrates data from 153 public and internal cohorts, covering multiple spatial transcriptomics technologies. The processing pipeline includes tissue segmentation, nuclei detection, and expression-morphology alignment, generating 2.1 million pairs. HEST-Benchmark evaluates on three tasks: pathology foundation models (e.g., UNI, CTransPath), biomarker discovery (e.g., gene expression prediction), and multimodal learning (image-text alignment). Limitations: data is biased towards cancer samples, with insufficient normal tissue coverage; batch effects need correction.
HEST-1k will drive computational pathology from narrow tasks towards general foundation models. For pharmaceutical companies, it can be used for target discovery and drug response prediction; for diagnostic companies, it can improve the generalization ability of AI pathology systems. The open-source nature of the dataset will promote industry-academia collaboration.
It is recommended that pathology AI companies use HEST-1k to pre-train or fine-tune models, improving performance on rare diseases and multiple cancer types. Investment institutions can focus on startups developing diagnostics or drug discovery based on this dataset.
Key points to watch: 1) Actual performance of HEST-1k in foundation model pre-training; 2) Expanded datasets contributed by the community; 3) Integration with clinical data; 4) Privacy and ethical issues (e.g., de-identification of patient data).