Research Experience

Ongoing Research and Key Contributions

  • Efficient Deployment and Inference of Large Language Models Research on optimizing the inference effectiveness, efficiency, and scalability of Large Language Models (LLMs) through strategies such as cloud-edge collaboration, distributed AI, and ML efficiency techniques (e.g., early exit, quantization).
    Mentor: Dr. Yanzhao Wu
    Contributions:
    • CE-CoLLM: Efficient and Adaptive Large Language Models Through Cloud-Edge Collaboration
      • arXiv:2411.02829, under review at MLSys 2025
        Proposed the CE-CoLLM method to optimize the inference efficiency and accuracy of LLMs on edge devices through cloud-edge collaboration and ML efficiency techniques, addressing diverse requirements such as inference accuracy, low latency, resource constraints, and privacy preservation.
    • DA-MoE: Dynamic Expert Allocation for Mixture-of-Experts Models
      • arXiv:2409.06669, 2024
        Proposed the DA-MoE method, a dynamic expert allocation mechanism for Mixture-of-Experts (MoE) models that leverages attention-based token importance in Transformer architectures to dynamically adjust the number of experts per token, enhancing efficiency and predictive performance.
  • Advanced Training Strategies and Ensemble Learning for Model Performance Research on enhancing training efficiency, performance, and robustness of deep neural networks (DNNs) and large language models (LLMs).
    Mentor: Dr. Yanzhao Wu
    Contributions:
    • Efficient and Learning Rate Boosted Deep Ensembles
      • under review at CVPR 2025
        Proposed the LREnsemble framework, effectively utilizing diverse models, generated through learning rate (LR) tuning, to construct efficient and high-quality ensembles, avoiding the waste of sub-optimal LR-tuned models by leveraging their diversity for ensemble learning.
    • Effective Diversity Optimizations for Deep Ensembles
      • CogMI 2024
        Proposed the Synergistic Diversity metric, significantly improving ensemble accuracy and robustness to out-of-distribution samples by optimizing diversity among member models.
    • Rethinking Learning Rate Tuning in Large Language Models
      • CogMI 2023
        Introduced the LRBench++, a dynamic learning rate tuning framework, improving DNNs and LLMs training efficiency and achieving a balance between model accuracy and training cost.

Academic Service and Peer Review Contributions

  • Reviewer: ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2024
  • Reviewer: Asian Conference on Machine Learning (ACML), 2024
  • External Reviewer: The Web Conference (WWW), 2024, 2025
  • External Reviewer: Association for the Advancement of Artificial Intelligence (AAAI), 2024
  • External Reviewer: International Joint Conference on Artificial Intelligence (IJCAI), 2024
  • External Reviewer: IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2024
  • External Reviewer: SIAM International Conference on Data Mining (SDM), 2024

Awards

  • IEEE TPS 2023 NSF Travel Award, November 2023