|
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
Sirnam Swetha,
Rohit Gupta,
Parth Parag Kulkarni,
David G Shatwell,
Jeffrey A Chan Santiago,
Nyle Siddiqui,
Joseph Fioresi,
Mubarak Shah
CVPR (Highlight), 2026
paper /
Data /
arxiv /
code /
project page /
bibtex
|
|
TIGeR: A Unified Framework for Time, Images and Geo-location Retrieval
David G Shatwell,
Sirnam Swetha,
Mubarak Shah
CVPR, 2026
paper /
supp /
arxiv /
project page /
bibtex
|
|
Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety
Younggun Kim*,
Sirnam Swetha*,
Fazil Kagdi,
Mubarak Shah
(*equally contributing first author)
CVPR, 2026
paper /
Data /
arxiv /
code /
project page /
bibtex
|
|
SMPRO: Self-Supervised Visual Preference Alignment via Differentiable Multi-Preference Multi-Group Ranking
Sirnam Swetha,
Rui Meng,
Shwetha Ram,
Tal Neiman,
Son Tran,
Mubarak Shah
AAAI, 2026
paper /
project page /
bibtex
|
|
GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space
David G Shatwell,
Ishan Rajendrakumar Dave,
Sirnam Swetha,
Mubarak Shah
ICCV (Oral), 2025
paper /
supp /
arxiv /
code /
project page /
bibtex
|
|
TimeLogic: A Temporal Logic Benchmark for Video QA
Sirnam Swetha,
Hilde Kuehne,
Mubarak Shah
arxiv, 2025
paper /
arxiv /
project page /
bibtex
|
|
The Telephone Game: Evaluating Semantic Drift in Unified Models
Sabbir Mollah,
Rohit Gupta*,
Sirnam Swetha*,
Qingyang Liu^,
Ahnaf Munir^,
Mubarak Shah
(*equal contribution)
arxiv, 2025
paper /
Data /
arxiv /
code /
project page /
bibtex
|
|
StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
Nyle Siddiqui,
Rohit Gupta,
Sirnam Swetha,
Mubarak Shah
arxiv, 2025
arxiv /
bibtex
|
|
SB-Bench: Stereotype Bias Benchmark for Large Multimodal Models
Vishal Narnaware*,
Ashmal Vayani*,
Rohit Gupta♠,
Sirnam Swetha♠,
Mubarak Shah
(*equally contributing first author, ♠equally contributing second author)
arxiv, 2025
paper /
Data /
arxiv /
code /
project page /
bibtex
|
|
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
Sirnam Swetha,
Jinyu Yang,
Tal Neiman,
Mamshad Nayeem Rizve,
Son Tran,
Benjamin Yao,
Trishul Chilimbi,
Mubarak Shah
ECCV, 2024
paper /
supplement /
arxiv /
project page /
bibtex
|
|
Multi-SK: Preserving Modality Structure Improves Multi-Modal Learning
Sirnam Swetha,
Mamshad Nayeem Rizve,
Nina Shvetsova,
Hilde Kuehne,
Mubarak Shah
ICCV, 2023
paper /
supplement /
arxiv /
code /
project page /
bibtex
|
|
Unsupervised Discriminative Embedding for Action Learning in Complex Activities
Sirnam Swetha,
Hilde Kuehne,
Yogesh S Rawat,
Mubarak Shah
ICIP, CVPR L2ID (Oral), 2021
paper /
arxiv /
project page /
bibtex
|
|
Sequence-to-Sequence Learning for Human Pose Correction in Videos
Sirnam Swetha,
Vineeth N Balasubramanian,
CV Jawahar
ACPR, 2017
paper /
project page /
bibtex
|
|
Efficient Object Annotation for Surveillance and Automotive Applications
Sirnam Swetha,
Anand Mishra,
Guruprasad M Hegde,
CV Jawahar
WACVW, 2016
paper /
project page /
bibtex
|
|
Online handwriting recognition using depth sensors
Rajat Aggarwal*,
Sirnam Swetha*,
Anoop M Namboodiri,
Jayanthi Sivaswamy,
CV Jawahar
(*equally contributing first author)
ICDAR, 2015
paper /
project page /
bibtex
|