Tourism destination image encompasses the multifaceted perceptions of visitors that influence their choice of destination. However, current computational methodologies often analyze visual and textual signals separately and seldom link image perception to the quality of recommendations. This study introduces multimodal destination image perception and recommendation (MM-DIPR), a framework utilizing public data that amalgamates visual encoding, natural language processing, geospatial context, and graph-based ranking into a cohesive perception-aware pipeline. MM-DIPR discerns seven interpretable perception attributes, namely, landmarkness, naturalness, cultural/historical salience, esthetic appeal, crowding proxy, sentiment affect, and activity affordance, through a dedicated perception bottleneck that facilitates cross-modal alignment and graph ranking. We establish a reproducible benchmark using solely public and openly licensed datasets (Google Landmarks Dataset v2, YFCC100M, Places365, Yelp Open Dataset, TREC Contextual Suggestion, Wikimedia/Wikidata/Wikivoyage, Overture Maps Places, and SNAP Gowalla/Brightkite) and assess MM-DIPR against 11 baselines (including spatiotemporal and knowledge-graph POI models) across six complementary tasks, namely, perception classification, image–text retrieval, personalized point-of-interest (POI) recommendation, cold-start recommendation, cross-city transfer, and explanation faithfulness. A formal spatial autocorrelation analysis confirms that all seven perception attributes exhibit statistically significant positive spatial autocorrelation (Moran’s I∈[0.289,0.421]; p<0.05), grounding the framework’s geospatial design choices. Experimental findings indicate that the perception bottleneck significantly enhances ranking accuracy (Recall@10 of 0.213 and NDCG@10 of 0.152 on the Yelp benchmark; p<0.01) and cold-start robustness compared to strong multimodal baselines. Ablation studies validate the independent contribution of each module, while diversity analysis reveals that calibrated re-ranking enhances long-tail exposure with minimal accuracy loss. A controlled annotation study with 2,000 images and three annotators confirms substantial inter-annotator agreement (κ∈[0.68,0.83]) for all attributes. Qualitative explanation panels ground recommendations in public visual and textual evidence, thus providing practical utility for destination management practitioners. All experimental code and derived benchmark artifacts will be made available to support reproducibility.