VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentation
Fabian Isensee, Constantin Ulrich, Yannick Kirchhoff, Klaus Maier-Hein, Maximilian Rokuss, Moritz Langenberg, Benjamin Hamm, Sebastian Regnery, Lukas Bauer, Efthimios Katsigiannopulos, Tobias Norajitra
cvpr
Research metadataShow detailsHide details
- Affiliations
- German Cancer Research Center, Division of Medical Image Computing, Germany · Faculty of Mathematicsand Computer Scienceand3Medical Faculty-Heidelberg University · Helmholtz Imaging, · Department of Radiation Oncology, Heidelberg University Hospital, Germany · HIDSS4Health, Heidelberg7Pattern Analysisand Learning Group, Heidelberg University Hospital · a) known prompts b) cross-modality c) unknown concept d) clinical language Real-World Dataset · Figure1. Vox Tellperforms3Dmedicalimagesegmentationdirectlyfromarbitraryfree-textprompts.Thefigureshowsprogressively · modalities,(c)novelconceptsneverencounteredduringtraining,and(d)clinicallanguageunderstandingfromrealradiologyreportswith · shownin(d),where Vox Telloutperformspriortext-promptablesegmentationmethods.
- Published
- 2026-06-01
- Processed
- 6/20/2026, 6:26:23 PM
- Analysis model
- Not available
- Analysis status
- not_analyzed
- Local PDF artifact
- papers/pdf/2026/voxtell-free-text-promptable-universal-3d-medical-image-segm.pdf
Abstract
We introduce VoxTell, a vision-language model for text-prompted volumetric medical image segmentation. It maps free-form descriptions, from single words to full clinical sentences, to 3D masks. Trained on 62K+ CT, MRI, and PET volumes spanning 1K+ anatomical and pathological classes, VoxTell uses multi-stage vision-language fusion across decoder layers to align textual and visual features at multiple scales. It achieves state-of-the-art zero-shot performance across modalities on unseen datasets, excelling on familiar concepts while generalizing to related unseen classes. Extensive experiments further demonstrate strong cross-modality transfer, robustness to linguistic variations and clinical language, as well as accurate instance-specific segmentation from real-world text. Code is available at: https://github.com/MIC-DKFZ/VoxTell
Stored markdown is an ingestion scaffold, so it remains excluded until substantive methodology, result, and limitation sections are available.
Analysis unavailable
Only an ingestion scaffold is stored; substantive analysis is not available yet.