Precise positioning and identification of unauthorized unmanned aerial vehicles (UAVs) are of crucial importance for spectrum security and privacy protection in future intelligent networks. Although various single-modality approaches have been investigated, their performance degrades under the sensor-specific noise, resulting in suboptimal performance and robustness. To address these security challenges, we propose a multi-layer radio frequency (RF)-vision fusion framework that synergistically exploits temporal-spectral features of UAV RF signals and spatial-visual information to achieve precise and robust UAV positioning and identification. Moreover, a corresponding unified RF-Vision fusion Network (RFViNet) is designed to exploit the RF-vision cross-modal complementary and semantic synergy. Specifically, by leveraging the novel RFinformed proposal generation, RF-enhanced feature modulation, and RF-guided semantic query modules, the RFViNet effectively exploits the complementary strengths of RF and visual modalities. Furthermore, a practical RF–vision platform is developed to evaluate the performance of our method under various challenging conditions. Experimental results on the real-world dataset demonstrate that the proposed method achieves a competitive 85.8% average precision AP50, highlighting its potential for enhancing the spectrum security in future intelligent wireless networks.