The School of EECS is hosting the following HDR Progress Review 3 Seminar:
 

Toward Efficient Lightweight Recommender Systems
 
Speaker: Jason Xurong Liang
 
Abstract:
Recommender systems have evolved rapidly from traditional ID-based latent factor models to advanced architectures leveraging Graph Neural Networks (GNNs) and Large Language Models (LLMs). While these advancements have drastically improved recommendation accuracy by capturing complex, high-order user-item interaction dynamics, they introduce critical deployment bottlenecks. The exponential increase in model complexity, like staggering embedding table storage costs, intensive graph propagation overhead and massive LLM parameter scales, prohibits the deployment of state-of-the-art recommenders on resource-constrained platforms like edge devices. This thesis addresses these critical hardware and computational barriers by proposing a series of highly efficient, lightweight architectures across all three foundational recommendation paradigms without compromising predictive effectiveness.
 
To address the parameter inefficiency inherent in traditional collaborative filtering, this thesis introduces Compositional Embeddings with Regularized Pruning (CERP). Existing compression techniques often struggle under strict storage constraints, either aggressively degrading embedding fidelity via pruning or inducing severe representational collisions via static hashing. CERP structurally bridges dynamic sparsification with parameter sharing by representing entities through combinations drawn from dual compact codebooks. Guided by a novel pruning regularizer, CERP encourages mutually complementary sparsification across these codebooks, maximizing embedding density and uniqueness under extremely stringent storage budgets.
 
To tackle both the massive storage and the severe runtime computation costs of GNN-based recommenders, we propose two sequential frameworks: Lightweight Embeddings for Graph Collaborative Filtering (LEGCF) and Lightweight Embeddings with Rewired Graph (LERG). LEGCF abandons pre-defined hash functions in favor of a learnable assignment matrix that dynamically captures semantic topology for meta-embedding composition. Building upon this foundation, LERG incorporates codebook quantization to further shrink the memory footprint and introduces an innovative graph rewiring mechanism to systematically prune low-contribution nodes. This server-edge deployment paradigm drastically reduces the multiplication and accumulation operations (MACs) required for iterative message passing, enabling seamless on-device fine-tuning and low-latency inference.
 
Navigating the inherent trade-off between model parameter size and representation capacity in LLM-based recommenders, we present the Fusion of Layer-wise Exits for Sequential Recommendation (FLEXRec). Scaling down LLMs to compact backbones severely weakens their semantic capacity, while common generative and reasoning workarounds introduce prohibitive autoregressive inference latency. FLEXRec resolves this by adopting a highly efficient discriminative retrieval paradigm that completely bypasses generative decoding. To compensate for the limited capacity of the compact backbone, FLEXRec utilizes an Adaptive Continuous Router (AC-Router) to dynamically fuse intermediate layer-wise exits based on individual user sequence complexity. Governed by a novel target-$k$ hinge loss, this multi-depth ensemble strategy gracefully controls computational sparsity, allocating deeper reasoning to complex interaction histories while maintaining peak efficiency for trivial sequences.
 
Through extensive experiments on massive, real-world benchmark datasets, the proposed frameworks consistently demonstrate superiority over state-of-the-art embedding optimization and lightweight baselines. By systematically resolving the architectural bottlenecks of traditional, graph-based, and language-based models, this thesis constructs a comprehensive pathway for deploying highly accurate, scalable, and adaptable recommender systems in resource-constrained environments.
 
Bio:
Jason Xurong Liang is a final-year Ph.D. student in Computer Science at the University of Queensland, Australia. He earned a Bachelor of Computer Science with First-Class Honors from the same university in 2022.
His Ph.D. research focuses on efficient lightweight recommender systems, advised by Assoc. Prof. Rocky Tong Chen and Prof. Hongzhi Yin. His broader interests include data mining, recommender systems, user modeling, and natural language processing.

About Data Science Seminar

This seminar series is hosted by EECS Data Science.

Venue

Room: 78 - 631/632 (MM Lab)
Zoom link: https://uqz.zoom.us/j/85662085365