Agentic Agricultural Assistant with Vision-Language Models
The School of EECS is hosting the following HDR Progress Review 1 Seminar:
Agentic Agricultural Assistant with Vision-Language Models
Speaker: Boyu Luo
Abstract:
Agricultural consultation is an important task for supporting crop health, plant disease management and farming decisions. In practice, farmers, growers and agronomists often provide field images and short text descriptions, expecting expert advice on plant diseases, pest damage, nutrient problems, and management actions. Recent vision-language models (VLMs) provide a promising foundation for building agricultural assistants because they can process multimodal inputs. Existing approaches have adapted multimodal models to agricultural knowledge and have built datasets and benchmarks via agricultural question answering, and expert-style response generation. These studies show that VLMs can support agricultural knowledge understanding and domain-specific response generation.
However, most current systems and evaluations still follow a direct question-answering format, where the model receives the available input and directly generates a final response. This format does not fully match real consultation, where the information may be incomplete and the final advice should be supported by evidence. This seminar presents two works that address these limitations. The first work, AgriTalk-RL trains VLMs to decide whether to ask for clarification or provide a final response. AgriTalk-RL uses reinforcement learning with gated consistency, self-correction, and LLM-as-a-judge rewards. The second work, AgriGym, reformulates agricultural consultation as an executable evidence-seeking process, where agents inspect images, retrieve related evidence, and submit evidence-grounded recommendations. Together, these works aim to move agricultural VLM systems beyond static response generation toward reliable agricultural consultation agents.
Bio: Boyu Luo is a PhD student at the School of Electrical Engineering and Computer Science, The University of Queensland. His research focuses on multimodal agentic reasoning of vision-language models. His current work studies how multimodal agents can make decisions under incomplete evidence and produce recommendations supported by reliable reasoning processes.
About Data Science Seminar
This seminar series is hosted by EECS Data Science.
Venue
Zoom: https://uqz.zoom.us/j/85865162268