Background: AI is increasingly being integrated into health care, making it important to understand stakeholder preferences for AI-enabled technologies. Although discrete choice experiments (DCEs) are widely used to elicit preferences, evidence on preference attributes, willingness to pay (WTP), and reporting quality in AI-related DCEs has not been systematically synthesized. Objective: This systematic review aims to synthesize stakeholder preferences for AI-enabled health care technologies elicited through DCEs and assess reporting quality using the DIRECT (DIscrete choice experiment REporting ChecklisT). The goal is to identify key preference attributes and methodological gaps to guide the development of AI technologies aligned with real-world needs. Methods: This systematic review was conducted in accordance with PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines. We systematically searched PubMed, Embase, Web of Science, Scopus, the Cochrane Library, and the International Health Technology Assessment (HTA) Database from database inception to March 2026. We included studies reporting original DCE data on AI-enabled health care technologies. Study characteristics, including experimental design features and econometric modeling approaches, were extracted and summarized descriptively. Attributes were systematically categorized using a structured framework informed by established HTA taxonomies. Stakeholder preference evidence was synthesized narratively. In addition, DCE design characteristics, preference outcomes, WTP estimates, and reporting quality indicators were extracted and synthesized. Reporting completeness was assessed using the DIRECT checklist. Results: Twenty-seven studies (28 DCEs) involving patients, clinicians, and the public were included, covering AI applications in diagnosis, screening, treatment, disease management, and decision support. Across the included studies, a total of 163 attributes were identified across 5 domains, and preference outcomes were synthesized at the level of 30 application contexts. Usability domain attributes were most frequently included, accounting for 36.81% (60/163) of all attributes, with presentation format being the most frequently reported usability attribute (27/60, 45.00%), but it was often rated least important (9/30, 30.00%). In contrast, performance attributes, especially effectiveness (17/30, 56.67%), were most often ranked as most important. Preference heterogeneity varied by context, with performance attributes dominating in clinical applications and usability attributes being more prominent in decision support. Conditional and mixed logit models were commonly used, while scale heterogeneity was rarely explored. Average reporting completeness was high (mean 84.47%, SD 11.49%), although design effects (10/27, 37.04%) and randomization (14/27, 51.85%) were inconsistently reported. Conclusions: DCE evidence suggests a mismatch between commonly included attributes and stakeholder priorities. Effectiveness is generally a key determinant of preferences, although preference structures are context dependent. Usability and performance attributes vary in importance across AI applications. Although overall reporting completeness was high, design effects and randomization were inconsistently reported. These findings highlight the need to align attribute selection with decision contexts and improve transparency in study design and reporting, which may enhance the interpretability of DCE evidence and support the development of AI technologies reflecting real-world health care needs.