XPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding
August 2, 2026 12 minutes
Welcome to the SSAIL Lab at UIUC. At SSAIL, we’re driven by a bold vision: to redefine how large-scale machine learning systems are built, optimized, and applied. On the systems side, we build efficient, effective, and easy-to-use systems that power the future of AI, from high-performance and distributed training to ultra-fast inference engines. On the machine learning side, we explore new frontiers, from model compression to vision-language models, agents, scientific AI, and beyond, pushing the boundaries of what’s possible while uncovering the principles that make it all work. What excites us most is the synergy between systems and algorithms. Every algorithmic insight drives new system breakthroughs, and every system we build opens the doors to new algorithmic capabilities. This virtuous cycle is where innovation thrives, and where SSAIL is helping shape the next generation of intelligent systems. Explore our latest blog posts to see what we’re working on.
August 2, 2026 12 minutes
This blog presents the motivation, design principles, and key results behind RecScale, a system-aware approach to scaling Deep Learning Recommendation Models (DLRMs) that addresses critical memory and communication bottlenecks in distributed training.
October 12, 2025 10 minutes
Efficient full-parameter fine-tuning of GPT-OSS-20B & Qwen3-14B models on a single NVIDIA GH200 and Llama3-70B on four NVIDIA GH200 Superchips, while delivering up to 600 TFLOPS training throughput.
October 7, 2025 5 minutes
This blog presents a deep analysis of Alpha-Fold 3 (AF3) training pipelines, pinpointing their inefficiencies and introduces MegaFold: an end-to-end training system for AF3 that addresses the aforementioned issues.
October 3, 2025 10 minutes
This blog presents the motivation, insights, and key optimizations behind VoltanaLLM, our system for energy-efficient LLM inference. We’ll walk through why energy matters, how conventional GPU frequency scaling falls short, the surprising behaviors we uncovered when profiling LLM serving, how P/D disaggregated serving creates unique opportunities, and how VoltanaLLM’s co-design of frequency control and routing achieves up to 36.3% GPU energy savings while maintaining near-perfect Service Level Objective (SLO) attainment.
September 14, 2025 9 minutes
This blog presents the background and key optimizations behind X-MoE, along with our hands-on experience scaling MoE model training on Frontier, the AMD GPU supercomputer.
August 24, 2025 10 minutes