Last updated in 2026/07

Xiangyun Ding

Education

University of California, Riverside

Ph.D student, Department of Computer Science

Co-advised by Prof. Yan Gu and Prof. Yihan Sun

Tsinghua University

Bachelor of Engineering, Department of Computer Science and Technology

Adviser: Prof. Wenjian Yu

Skills

Languages: C++, Python, Go, Java

AI/LLM: PyTorch, CUDA/ROCm programming, Triton/Tilelang

Awards and Honors

Work Experience

Software Engineer Intern

CasualFlow Inc.

  • Optimize the LLVM AMDGPU backend. The new instruction scheduling algorithm gives up to 1.2x performance gain.
  • Develop high-performance GPU kernels for LLM inference using CUDA/Triton/Tilelang, including GEMM, FlashAttention/KDA, Fused MoE/MegaMoE, in bf16/fp8/mxfp4 weights.
  • Develop In-Context Reinforcement Learning (ICRL) for LLM-driven program optimization.

Software Engineer

Airbnb, Inc.

  • Host tiering project backend development. Process data with airflow. Develop Java backend services.
  • Airbnb chat-bot development. Build new bot rules for user-bot interaction. Use ML models to direct bot behaviors.
  • Coding Interviewer.

Software Engineer Intern

Google

  • Processed and analyzed large-scale Wikipedia revision histories for the WikiTrust vandalism detection project.
  • Designed and implemented rule-based feature extraction and heuristic classifiers.

Publications

Preprints

  1. 2026
    ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants πŸ”—

    Haohui Mai, Xiaoyan Guo, Xiangyun Ding, Daifeng Li, Qiuchu Yu, Chenzhun Guo, Cong Wang, Jiacheng Zhao, Christos Kozyrakis, Binhang Yuan.

    Accepted for ACM SIGOPS Symposium on Operating Systems Principles (SOSP), 2026.

Publications

  1. 2026
    Parallel Metric Skip Lists and Nearest Neighbor Search πŸ”—

    Xiangyun Ding, Rohin Garg, Yan Gu, Yihan Sun. (Authors listed alphabetically)

    In 38th ACM Symposium on Parallelism in Algorithms and Architectures (SPAA), 2026.

  2. 2025
    New Algorithms for Incremental Minimum Spanning Trees and Temporal Graph Applications πŸ”—

    Xiangyun Ding, Yan Gu, Yihan Sun.

    In SIAM Conference on Applied and Computational Discrete Algorithms (ACDA), 2025.

  3. 2024
    Parallel and (Nearly) Work-Efficient Dynamic Programming πŸ”—

    Xiangyun Ding, Yan Gu, Yihan Sun.

    In 36th ACM Symposium on Parallelism in Algorithms and Architectures (SPAA), 2024. Outstanding paper award.

  4. 2024
    Fast and Space-Efficient Parallel Algorithms for Influence Maximization πŸ”—

    Letong Wang, Xiangyun Ding, Yan Gu, Yihan Sun.

    In Proceedings of the VLDB Endowment (VLDB), 2024.

  5. 2023
    Efficient Parallel Output-Sensitive Edit Distance πŸ”—

    Xiangyun Ding, Xiaojun Dong, Yan Gu, Youzhe Liu, Yihan Sun. (Authors listed alphabetically)

    In 31st Annual European Symposium on Algorithms (ESA), 2023. Best paper award.

  6. 2022
    DP-Nets: Dynamic programming assisted quantization schemes for DNN compression and acceleration πŸ”—

    Dingcheng Yang, Wenjian Yu, Xiangyun Ding, Ao Zhou, Xiaoyi Wang.

    In Integration, the VLSI Journal, 2022.

  7. 2020
    Efficient model-based collaborative filtering with fast adaptive PCA πŸ”—

    Xiangyun Ding, Wenjian Yu, Yuyang Xie, Shenghua Liu

    In IEEE 32nd International Conference on Tools with Artificial Intelligence (ICTAI), 2020.

Other Projects

LQ-Nets-PyTorch πŸ”—

  • A PyTorch implementation of LQ-Nets (neural network quantization).

Parallel-GPT-2 πŸ”—

  • A C++ implemented GPT-2, running on multi-CPUs in parallel.
  • Up to 3x faster than llm.c on 64 cores.

Vahadane: Histological Image Normalization πŸ”— (80+ stars)

Single-user Relational Database πŸ”—

Five-level Pipelined CPU πŸ”—

  • A five-level pipelined CPU implementation on FPGA using Verilog and the MIPS instruction set.

Distributed Sharded Key/Value Storage πŸ”—

  • A distributed KV database written in Go. Implemented the Raft consensus algorithm.