Inclusive Multimodal Understanding for People with Disabilities
The School of EECS is hosting the following HDR Progress Review 1 Confirmation Seminar:
Inclusive Multimodal Understanding for People with Disabilities
Speaker: Yan Ke
Abstract: Inclusive multimodal understanding is increasingly relying on Multimodal Large Language Models (MLLMs), yet current systems remain fundamentally limited in disability-centered scenarios, where people with limb deficiencies are rarely represented in mainstream vision-language datasets. Existing vision-language benchmarks mainly evaluate general object, scene, and action understanding, while inclusive VQA datasets mostly focus on blind and low-vision users, failing to systematically assess disability-specific visual reasoning. To address these gaps, this research introduces IVQA-LD, a large-scale expert-annotated benchmark for limb-deficiency-aware multimodal understanding, covering perception, classification, reasoning, and generation tasks. It further proposes Body-centric Structure-aware Initialization (BSI) to guide VLMs toward relevant limb-structural cues during fine-tuning. Future objectives focus on extending this research to real-world video understanding by developing models and benchmarks for evidence-aware multimodal reasoning in disability-centered scenarios, ultimately contributing to more reliable human-centered multimodal AI systems.
Bio: Yan Ke is a Ph.D. student at the School of EECS, University of Queensland, under the supervision of AsPr Xin Yu. Her research centres on human-centric understanding.
About Data Science Seminar
This seminar series is hosted by EECS Data Science.