VISHAM-KG is a multimodal framework that constructs knowledge graphs from Hindi visual documents by aligning textual and visual entities. It combines rule-based linguistic analysis with computer vision techniques to produce subject-relation-object triplets in low-resource Indic language settings.