LLM agents and lossy compression benchmark teaser

LLM Agents Meet Lossy Compression: Benchmark, Demystify and Optimize across HPC Architectures

Changqing Li, Sheng Di, Kai Zhao, Wenqian Dong

ACM/IEEE International Conference for High Performance Computing, Networking, Storage, and Analysis (SC'26)

High-Performance ComputingLLM Systems
Reliable machine-learning surrogate workflow

Towards Building Reliable Machine Learning Surrogates for Scientific Applications

Bohan Zhang, Wenqian Dong, Guanpeng Li

IEEE Cloud Summit 2026

AI for ScienceMachine Learning Systems
Flying Serving system overview

FLYING SERVING: On-the-Fly Parallelism Switching for Large Language Model Serving

Best Paper Award nominee

Shouwei Gao, Junqi Yin, Feiyi Wang, Wenqian Dong

ACM International Conference on Supercomputing (ICS'26)

LLM SystemsMachine Learning Systems
LUMOS scientific machine-learning workflow

LUMOS: Democratizing SciML Workflows with L0-Regularized Learning for Unified Feature and Parameter Adaptation

Shouwei Gao, Xu Zheng, Dongsheng Luo, Sheng Di, Wenqian Dong

IEEE International Parallel & Distributed Processing Symposium (IPDPS'26)

AI for ScienceScientific Machine Learning
Evaluation of HPC coding agents

Evaluating LLM Coding Agents on SZ-Family Lossy Compression Across Architectures

Changqing Li, Shouwei Gao, Kai Zhao, Sheng Di, Wenqian Dong

IPDPS HPAI4S'26 WorkshopIEEE IPDPS Workshops

High-Performance ComputingLLM Systems
Phoenix wafer-scale engine architecture

Enabling Unstructured Sparse Fine-Tuning and Inference for Foundation Models on Wafer-Scale Engine

Haoyu Zheng, Yifan Zeng, Linghao Song, Murali Emani, Wenqian Dong

SC'25 ExHetAI Workshop

AI AcceleratorsMachine Learning Systems
HurriCast tropical cyclone generation poster

HurriCast: Synthetic Tropical Cyclone Track Generation for Hurricane Forecasting

Shouwei Gao, Meiyan Gao, Yuepeng Li, Wenqian Dong

AAAI 2025 Spring Symposium Series

AI for ScienceGenerative AI
Hybrid simulation and AI metadata workflow

Framework for tracking metadata, lineage and model provenance in hybrid simulation-AI HPC exascale workflows

Martin Foltin, Andrew Shao, Rishabh Sharma, Shreyas Kulkarni, Annmary Justine Koomthanam, Aalap Tripathy, Cong Xu, Wenqian Dong, Suparna Bhattacharya, Brian Sammuli, Paolo Faraboschi

CUG'25: Proceedings of the Cray User Group

AI for ScienceHigh-Performance Computing
TimeX++ time-series explanation framework

TimeX++: Learning Time-Series Explanations with Information Bottleneck

Zichuan Liu, Tianchun Wang, Jimeng Shi, Xu Zheng, Zhuomin Chen, Lei Song, Wenqian Dong, Jayantha Obeysekera, Farhad Shirani, Dongsheng Luo

41st International Conference on Machine Learning (ICML'24)

Explainable AITime Series
Auto-HPCnet framework

Auto-HPCnet: An Automatic Framework to Build Neural Network-based Surrogate Models for HPC Applications

Wenqian Dong, Gokcen Kestor, Dong Li

ACM International Symposium on High-Performance Parallel and Distributed Computing (HPDC'23)

AI for ScienceHigh-Performance Computing
Betty graph partitioning system

Betty: Enabling Large-Scale GNN Training with Batch-Level Graph Partitioning

Shuangyan Yang, Minjia Zhang, Wenqian Dong, Dong Li

International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS'23)

Graph Neural NetworksMachine Learning Systems
Fauce cardinality-estimation method

Fauce: Fast and Accurate Deep Ensembles with Uncertainty for Cardinality Estimation

Jie Liu, Wenqian Dong, Qingqing Zhou, Dong Li

International Conference on Very Large Data Bases (VLDB'21)

Database SystemsMachine Learning
MD-HM molecular dynamics system

MD-HM: Memoization-based Molecular Dynamics Simulations on Big Memory System

Zhen Xie, Wenqian Dong, Jie Liu, Ivy Peng, Yanbao Ma, Dong Li

ACM International Conference on Supercomputing (ICS'21)

High-Performance ComputingScientific Computing
Tahoe GPU inference engine

Tahoe: Tree Structure-Aware High Performance Inference Engine for Decision Tree Ensemble on GPU

Zhen Xie, Wenqian Dong, Jiawen Liu, Hang Liu, Dong Li

ACM European Conference on Computer Systems (EuroSys'21)

GPU ComputingMachine Learning Systems
Smart-PGSim power-grid simulation workflow

Smart-PGSim: Using Neural Network to Accelerate AC-OPF Power Grid Simulation

Wenqian Dong, Zhen Xie, Gokcen Kestor, Dong Li

ACM/IEEE International Conference for High Performance Computing (SC'20)

AI for ScienceHigh-Performance Computing
Adaptive neural-network fluid simulation

Adaptive Neural Network-Based Approximation to Accelerate Eulerian Fluid Simulation

Wenqian Dong, Jie Liu, Zhen Xie, Dong Li

ACM/IEEE International Conference for High Performance Computing (SC'19)

AI for ScienceHigh-Performance Computing
Application resilience model

Modeling Application Resilience in Large Scale Parallel Execution

Kai Wu, Wenqian Dong, Qiang Guan, Nathan DeBardeleben, Dong Li

International Conference on Parallel Processing (ICPP'18)

High-Performance ComputingResilience