Transformer-based architectures have become central to medical image analysis, yet their practical value remains difficult to assess because studies vary widely in tasks, datasets, validation protocols, baselines, and reporting quality. This survey critically reviews recent transformer-based, hybrid, foundation, and transformer-alternative models across segmentation, classification, reconstruction, and image registration. A total of 128 studies published between 2021 and 2026 are organized using a task-, modality-, and architecture-aware taxonomy, with reported performance synthesized alongside baseline comparisons, reproducibility, computational cost, and clinical-readiness evidence. The findings indicate that the most convincing gains arise from task-adapted hybrid designs that combine local feature extraction with global context modeling, rather than from an unconditional superiority of transformers over convolutional networks. Persistent gaps include non-standardized benchmarks, limited external validation, incomplete code and weight availability, inconsistent efficiency reporting, weak uncertainty analysis, and insufficient clinical evaluation. Progress will require transparent reporting, multicenter validation, and clinically grounded assessment.