A multimodal cross-modal explainable retrieval-augmented generation framework for hallucination-grounded visual question answering

Babasaheb Satpute et al.