Background: AI and multimodal fusion are shifting orthodontics from experience-based to data-driven paradigms, but a rigorous systematic synthesis is lacking.
Objective: To conduct a PRISMA 2020 systematic review summarizing current applications, performance, and challenges of AI in orthodontic image analysis, intelligent workflows, and multimodal diagnostic integration.
Methods: Searches in PubMed, IEEE Xplore, Web of Science, and CNKI (2018–2026). Inclusion: original research, reviews, or technical reports applying AI or multimodal fusion to orthodontics; excluding conference abstracts, editorials, non-English/Chinese. Two reviewers independently screened and extracted data. QUADAS-2 assessed bias for diagnostic accuracy studies. Narrative synthesis.
Results: Thirty-one representative studies (including two systematic reviews) were included. In image analysis, deep learning-based automated cephalometric landmark detection (mean radial error 0.8–1.2 mm for 2D landmarks, 87% acceptance within 2 mm), 3D CBCT segmentation (tooth Dice >0.94), and cervical vertebral bone age assessment (accuracy 95%) have achieved levels approaching or comparable to clinical experts. Within intelligent workflows, automated tooth arrangement (reinforcement learning reducing round-tripping by 30%), bracket positioning (PointNet accuracy 98.93%, 2.9 ms per tooth), and full-cycle management systems including chatbot-based remote monitoring have markedly enhanced efficiency and precision. Multimodal fusion frameworks integrating imaging, intraoral scanning, clinical text, and biomechanical data (e.g., DeepMSM for midpalatal suture assessment with 93.75% accuracy) demonstrated diagnostic capabilities beyond unimodal approaches. Persistent challenges include data heterogeneity, limited interpretability, and lack of external clinical validation.
Conclusions: Current evidence indicates that AI and multimodal fusion can substantially improve orthodontic efficacy, but most models remain at the experimental stage. More high-quality studies adhering to PRISMA standards and real-world validation are urgently needed. Emerging frontiers such as digital twins, multimodal foundation models, and embodied intelligence warrant further investigation.