MUG'26
(Preliminary Program)
All Times Are U.S. EDT
Bale Theater at Ohio Supercomputer Center
Monday, August 17, 2026
Abstract
As HPC and AI converge, the fabric must serve both latency-sensitive MPI traffic and bandwidth-hungry AI training. This talk presents the Cornelis Networks CN7000 distributed compute fabric — a 1.6 Tb/s-per-port, 72-port switch paired with a PCIe Gen7 NIC — converging Ultra Ethernet, UALink, RoCEv2, and standard Ethernet on one fabric. We cover the features that matter most to this community: hardware message matching, multi-path reliable transport, end-to-end congestion management with in-band signaling and link-level retry, and in-network collectives via programmable RISC-V vector engines at every switch port. On the software side, we walk through the libfabric provider path, verbs, and NCCL-over-OFI, and introduce EarlySim — our datamodel-driven simulation environment that lets MPI developers, including the MVAPICH community, co-design and validate against CN7000 ahead of silicon. We close with a preview of CN8000, our next-generation fabric: multiplied bandwidth with a fully distributed, scheduler-less switching architecture.
Bio
As CTO of Cornelis Networks, Charles oversees the company’s hardware and software technology stack, including architectural direction, roadmaps, and innovation. Charles has previously held senior leadership roles at IBM, Intel, Akuna Capital, and Jump Trading. At IBM, Charles managed the design and deployment of job launch and message-passing architecture for Blue Gene, Power, and x86 systems. At Intel, he advanced MPI implementations for high-performance networks and led the development and open-sourcing of libfabric. At Cornelis Networks, Charles initiated innovative software solutions for high-performance interconnects. At Akuna Capital and Jump Trading, he led teams developing ultra-low latency trading platforms. Charles holds master’s degrees in chemistry from Columbia University and computer science from the University of Minnesota.
Abstract
SRAM-based accelerators are turning low-latency inference into a practical reality. As models outgrow any single accelerator, the dominant question shifts from the chip to the system: how a model is partitioned and how the resulting pieces are connected without sacrificing the latency advantage of the underlying hardware. The interfaces and protocols built for prior generations of accelerators do not scale to the demands of modern inference—wrapping a highly capable accelerator in legacy interconnect quickly renders that interconnect the bottleneck.
This tutorial presents the communication model developed at d-Matrix to address this problem, one that remains consistent even as the underlying transport changes—whether a message crosses a chiplet, a package, a server, or a rack. It describes how streaming-mode communication across nodes over Ethernet, homogeneous or heterogeneous, underpins the disaggregation techniques shaping modern ML inference. This is enabled by two components: JetStream, a network interface for low-latency inference clusters built around Corsair, and Aviator Fabric, a control-plane software layer that makes transport transparent at scale.
To support these communication schemes, we developed a hardware-software co-designed approach to device-initiated communication that enables the lowest-latency inference path. JetStream extends interconnectivity in the inference regime into novel heterogeneous disaggregation, improving both throughput efficiency (Throughput/kW) and interactivity (TPS/user).
Bios
Sai Rahul Chalamalasetti is a Sr. Principal AI Systems Architect at d-Matrix, where he serves as the Architect for the JetStream IO Accelerator card, the SoC IO-Block Architect (Die-to-Die and Scale-Up) for the next-generation transformer accelerator, and the lead for Scale-Out connectivity for low-latency inference. Previously, he worked at Hewlett Packard Labs as a Researcher and at HPE Servers as a Hardware Engineer.
A Senior Member of IEEE, he holds 14 approved patents with 10 additional USPTO applications accepted and has authored over 30 conference and workshop publications. His research interests include datacenter systems architecture, heterogeneous architectures, interconnect fabrics, and FPGAs.
He received his Ph.D. and M.S. in Computer Engineering from the University of Massachusetts Lowell in 2012 and 2009, respectively.
David Karlov leads JetStream software development at d-Matrix, building next-generation AI inference infrastructure. Over more than twenty years, he has combined hands-on engineering with technical leadership in the development of high-performance software systems spanning graphics, computer vision, machine learning and distributed systems, and is an inventor on numerous patents in imaging and graphics technologies. His experience has focused on extracting performance from complex hardware and software systems, from rendering pipelines to modern AI infrastructure.
10:00 - 10:30
Coffee Break, Posters and Demos
Abstract
Single-collective benchmarks give us a clean, peak snapshot of AI fabric performance, but they fail to reflect how real training workloads behave over time. In practice, AI training is inherently iterative, where each collective inherits congestion, queue state, and load-balancing decisions from previous steps. As a result, the key questions shift from peak throughput to whether the network converges, remains stable, and avoids long-tail latency under sustained load. By moving to iteration-aware and epoch-aware testing, we can expose hidden behaviors like congestion buildup, oscillation, and periodic resets at epoch boundaries.
Bios
Ankur Sheth is Senior Director for Strategic Projects at Keysight Technologies. In his current role he oversees the Network Test group’s AI initiatives including KAI DC Builder. With more than two decades of experience in the networking industry, he brings a unique perspective shaped by his expertise across engineering, product management and product marketing. His passion for networking drives him to create breakthrough products for new markets with a strong commitment to placing customer needs at the center of every decision.
Ankur holds a Masters’ in Business Administration from Indian Institute of Management, Ahmedabad and a Masters in Science degree in Electrical Engineering (Computer Networks) from University of Southern California, CA.
Eric Yu is a Solutions Architect focused on AI infrastructure validation and high-performance networking. His expertise spans AI fabrics, collective communication, RoCEv2, SONiC, congestion control, and large-scale distributed systems. He develops testing methodologies and workload emulation solutions that help organizations build reliable and scalable AI networks.
Abstract
The talk will present an HPC-AI software ecosystem of tools available on AWS, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and IBM Cloud based on MVAPICH MPI with support for GPUs. ParaTools Pro for E4S(TM) is a cloud image that includes agentic AI tools such as NVIDIA NemoClaw, OpenClaw, Google ADK and other AI tools such as NVIDIA NeMo(TM) and NVIDIA BioNemo(TM) optimized for GPUs and integrates with a performant remote-desktop based on Adaptive Computing’s Heidi AI/ODDC. It includes MVAPICH MPI, and supports both SLURM batch scheduler. On AWS, it supports x86, aarch64, and Trainium and Inferential nodes. E4S is a curated, Spack based software distribution of 100+ HPC, EDA, and AI/ML packages. It features AI tools such as TensorFlow, PyTorch, NVIDIA NeMo, NVIDIA BioNeMo, Ollama, vllm, Huggingface CLI, JAX, OpenAI, Google's Gemini API based chatbot, and other supporting tools including langchain, langgraph, pandas, and SciKit-Learn and supports AWS’ EFA, Google's IPUs, Infiniband on Azure, Oracle Mellanox Connect-X, IBM NVNICs with the optimized runtimes including MVAPICH MPI, NVIDIA NCCL, CUDA, NVIDIA Inference Xfer Library (NIXL), and NVSHMEM. It includes Codium, an IDE, Jupyter notebook, and visualization tools such as VisIt and ParaView all launched from a web browser without installing any additional software. This multi-user, multi-node, multi-gpu cloud image uses E4S and Spack as the core components for product integration and deployment of a range of HPC and AI/ML tools. These include performance evaluation tools such as TAU, HPCToolkit, DyninstAPI, PAPI, etc. and support both bare-metal and containerized deployment for CPU and GPU platforms. Container runtimes featured in the image include Docker, Podman, Singularity, and Charliecloud. Both Ubuntu 26.04 LTS and Rocky Linux 9.7 versions of Linux are supported. E4S is a community effort to provide open-source software packages for developing, deploying, and running scientific applications and tools on HPC platforms. It has built a comprehensive, extensible, coherent software stack that enables application developers to productively develop highly parallel applications that effectively target commercial cloud platforms.
Bio
Sameer Shende serves as a Research Professor and the Director of the Performance Research Laboratory at the University of Oregon and the President and Director of ParaTools, Inc. He serves as the technical lead of the E4S, ParaTools Pro for E4S(TM), TAU Performance System(R), and Program Database Toolkit (PDT) projects. His research interests include scientific software stacks, performance instrumentation, compiler optimizations, measurement, and analysis tools for HPC. He received his B.Tech. in Electrical Engineering from IIT Bombay in 1991, and his M.S. and Ph.D. in Computer and Information Science from the University of Oregon in 1996 and 2001 respectively.
12:00 - 1:30
Lunch Break, Posters and Demos (Cont'd)
Abstract
AI is ready to leave the lab, but the network is not ready to carry it. AI infrastructure is shifting from a compute-bound era to a connectivity-bound era. GPUs, accelerators, memory systems, and model architecture have scaled rapidly, but the network can no longer remain a passive utility. As AI workloads move into production-scale distributed deployments, interconnect bandwidth, latency, jitter, path control, observability, and long-distance throughput determine whether expensive compute capacity is fully utilized or remains stranded across hyperscalers, NeoClouds, regional data centers, and AI edge locations.
This presentation positions Scale Across as the next evolution beyond traditional data center interconnect. While conventional DCI was built for disaster recovery, storage replication, and workload migration, AI Scale Across extends backend networks for training and frontend networks for inference while preserving performance across distance. For training, the network gates utilization, collective communication, asynchronous training, and selected model-parallel approaches. For inference, distributed capacity depends on latency-aware routing, geo-distribution, KV-cache architecture, and predictable time-to-first-token performance
The proposed architecture introduces Fluid Interconnect: a programmable, xAware fabric that is traffic-aware, distance-aware, latency-aware, jitter-aware, network-aware, performance-aware, and tier-aware. This presentation introduces the 1Finity optical SmartNIC as a critical component of low-latency, high-throughput Scale Across. By placing transport acceleration at the server endpoint, the architecture combines a simplified data path, Verbs-style access, adaptable FPGA acceleration, and long-reach optical connectivity, while complementing existing intra-data-center fabrics
The objective is to convert remote, power-constrained, or stranded AI resources into usable computes through optical reach, intelligent endpoint acceleration, deterministic transport, and operational simplicity. This creates a scalable roadmap for distributed AI training, AI inference, HPC, remote GPU processing, and data-intensive platforms.
Bio
Dr. Solyman Ashrafi serves as Chief Technology Strategist at 1Finity/Fujitsu and is an accomplished executive leader with more than three decades of experience driving strategic value creation across Telecommunications, Hyperscale Cloud, and Artificial Intelligence. His career reflects a distinctive blend of technology vision, commercial execution, and enterprise-scale leadership, with a consistent focus on architecting next-generation ecosystems and translating advanced research and development into measurable business outcomes. Dr. Ashrafi advises senior executives, partners, and industry stakeholders on the economics and architecture of AI infrastructure, with particular emphasis on addressing stranded computes through the strategic alignment of interconnect, computing, sovereignty, and workload placement. His leadership spans hyperscaler, neutral colocation, and neoscalers interconnect architecture; 5G and edge network evolution; secure quantum key distribution; fluid optical interconnects; and agentic orchestration models that dynamically optimize paths, wavelengths, and bandwidth in response to policy, reliability, and energy requirements. Throughout his career, he has held multiple CTO and senior executive roles across global operators and leading technology vendors, including Nortel and Ericsson, where he converted long-range technology roadmaps into scaled commercial deployments. Previously, he played a key role in scaling metroPCS and supported its $33 billion merger with T-Mobile through product strategy, market execution, and business transformation. In addition to his operating leadership, Dr. Ashrafi has served as Managing Partner at a private equity firm, founded technology and software ventures, and contributed through advisory and board roles. He holds a PhD and advanced degrees in engineering and science and is the inventor of 140 patents and author of 85 refereed publications.
Abstract
As AI infrastructure evolves from isolated accelerator clusters to globally distributed systems, networking has become a first-order determinant of performance, efficiency, and scalability. This tutorial explores how Ethernet fabrics are emerging as the foundational interconnect technology for modern AI systems across scale-up, scale-out, and scale-across domains. The presentation will examine the differing communication patterns and performance requirements of tightly coupled GPU/XPU scale-up fabrics, large-scale distributed training and inference clusters, and region-scale AI deployments spanning multiple data centers. Topics include low-latency Ethernet switching, lossless transport mechanisms, congestion management, memory-semantic communication models, and the role of open standards such as Ultra Ethernet and emerging ESUN initiatives. Audience will gain an understanding of how advances in Ethernet switching silicon, transport protocols, and AI-aware networking architectures are enabling high-performance AI infrastructure at unprecedented scale. The tutorial will also discuss practical system architectures used in contemporary AI clusters
Bio
Niranjan Vaidya is a Distinguished Engineer and Architect in Broadcom's Core Switching Group (CSG). He leads switch and system architecture for the Tomahawk, Tomahawk Ultra and Trident families of high-performance Ethernet switches.
His focus areas span scale, performance and resiliency for compute fabrics. His contributions include high-performance programmable packet processing architectures, adaptive routing and congestion control, silicon friendly streaming algorithms, distributed synchronization at packet granularities, high-performance network telemetry, and an alphabet soup of networking protocols.
He is actively involved in the Ultra Ethernet Consortium and is the primary author of the Unified Forwarding Header (UFH) specification.
Niranjan has an MS in Computer Science from Iowa State University, and a B.Tech. with Honors in Computer Engineering from VJTI, University of Mumbai.
3:00 - 3:30
Coffee Break, Posters and Demos (Cont'd)
Abstract
The MVAPICH project has been powering HPC systems around the world for 25 years. This tutorial will provide an overview of the MVAPICH and MVAPICH-Plus libraries, the OSU MicroBenchmark (OMB) suite, and their features. We will focus on installation guidelines, runtime optimizations, and tuning parameters of these libraries with both microbenchmark and application use cases, as well as provide details on some of the latest features of MVAPICH.
Bios
Nat Shineman is a software engineer in the Department of Computer Science and Engineering at the Ohio State University. His current development work includes high performance interconnects, parallel computing, scalable startup mechanisms, and performance analysis and debugging of the MVAPICH2 library.
Benjamin Michalowicz is a PhD student at the Ohio State University under Prof. DK Panda and Prof. Hari Subramoni in the Network-Based Computing Laboratory. His research interests lie include high-performance computing (HPC), parallel/computer architectures, network-based computing for HPC, security in HPC, and parallel programming environments. Specifically, he is interested in efficiently offloading parallel programming models and computational workloads to Smart Network Cards like NVIDIA's BlueField DPUs. Ben actively contributes to the MVAPICH software and is a student member of the ACM and IEEE. Contact him at michalowicz.2@osu.edu
Abstract
Modern AI training and inference increasingly rely on high-performance communication and scalable software stacks to efficiently utilize modern HPC systems. This tutorial presents an overview of the HPC-AI software stack, a vendor-neutral, MPI-driven ecosystem that integrates leading AI frameworks, including PyTorch, TorchTitan, vLLM, and SGLang, with the MVAPICH-Plus communication runtime to enable high-performance distributed training and inference across modern CPU/GPU clusters. The tutorial covers distributed training and inference workflows, communication and runtime optimizations, and support for heterogeneous hardware and high-performance interconnects. It also discusses emerging trends in scalable AI systems and practical considerations for deploying and optimizing large AI workloads on HPC platforms.
Bios
Nawras Alnaasan is a Graduate Research Associate at the Network-Based Computing Laboratory, Columbus, OH, USA. He is currently pursuing a Ph.D. degree in computer science and engineering at The Ohio State University. His research interests lie at the intersection of deep learning and high-performance computing. He works on advanced parallelization techniques to accelerate the training of Deep Neural Networks and exploit underutilized HPC resources covering a wide range of DL applications including supervised learning, semi-supervised learning, and hyperparameter optimization. He is actively involved in several research projects including HiDL (High-performance Deep Learning) and ICICLE (Intelligent Cyberinfrastructure with Computational Learning in the Environment). Alnaasan received his B.S. degree in computer science and engineering from The Ohio State University. Contact him at alnaasan.1@osu.edu.
Jinghan Yao is a Graduate Research Associate at The Ohio State University, specializing in high-performance computing (HPC) and large-scale communication optimization for AI training and inference. His research develops efficient communication runtimes and libraries aimed at accelerating foundation model workloads. Jinghan has published papers at leading conferences including MLSys, IPDPS, and NeurIPS, and has presented his work at NVIDIA GTC 2024 and 2025. He collaborates closely with industry research teams such as Microsoft DeepSpeed to advance scalable communication techniques for AI.
5:00 - 6:30
Visit to the State of Ohio Computer Center, SOCC (Optional)
6:30
Reception and Dinner at Endeavor Brewing and Spirits
Tuesday, August 18, 2026
8:30 - 8:35
Opening Remarks
David Hudak, Ohio Supercomputer Center and Dhabaleswar K (DK) Panda, The Ohio State University
Abstract
To Be Determined
Bio
Dr. Dan Stanzione, Associate Vice President for Research at The University of Texas at Austin and Executive Director of the Texas Advanced Computing Center (TACC), is a nationally recognized leader in high performance computing, and has been involved in supercomputing for more than 30 years. He is the principal investigator (PI) for a number of the National Science Foundation (NSF) supercomputers, including the current Frontera system, which is the fastest supercomputer at a U.S. university, and is leading the upcoming NSF Leadership Class Computing Facility. Stanzione received his bachelor's degree in electrical engineering and his master's degree and doctorate in computer engineering from Clemson University.
Abstract
To Be Determined
Bio
DK Panda is a Distinguished Professor of Engineering and University Distinguished Scholar at the Ohio State University. He has published over 500 papers in the area of high-end computing and networking. The MVAPICH (High-Performance MPI over InfiniBand, iWARP, RoCE, Omni-Path, EFA, Rockport Networks, and Slingshot) libraries, designed and developed by his research group (mvapich.cse.ohio-state.edu), are currently being used by more than 3,450 organizations worldwide (in 92 countries). More than 1.92 million downloads of this software have taken place from the project's site. This software is empowering several InfiniBand clusters (including the 21st, 67th, and 88th ranked ones) in the TOP500 list. High-performance and scalable solutions for deep learning and machine learning from his group are available from hidl.cse.ohio-state.edu. High-performance and scalable libraries for Big Data stacks (Spark, Hadoop, and Memcached) and Data science applications from his group (hibd.cse.ohio-state.edu) are also publicly available. These libraries are currently being used by more than 370 organizations in 39 countries. More than 51,000 downloads of these libraries have taken place. He is a Fellow of ACM and IEEE, a recipient of 2022 IEEE Charles Babbage Award, and a recipient of the 2024 IEEE TCPP Outstanding Service and Contributions Award. More details about Prof. Panda are available at cse.ohio-state.edu/~panda.
10:15 - 10:45
Coffee Break, Posters and Demos (Cont'd)
Abstract
As distributed model training scales to span hundreds of thousands of GPUs, scale-out networks face unprecedented performance and efficiency demands. NVIDIA Spectrum-X Ethernet has been designed from the ground up to achieve predictable and stable network performance with high utilization and low latency. I will present the Spectrum-X multiplane architecture, which replaces hierarchical depth with topological parallelism, and introduces hardware-accelerated load balancing in NICs and switches as the key architectural approach to provide fast reaction to highly dynamic network conditions at the microsecond timescales that AI training workloads demand. I will describe the motivation, design principles, evaluation methodology and performance on state-of-the-art benchmarks, as well as the lessons we learned from deploying and debugging Spectrum-X networks in large-scale systems. The evaluation highlights production-grade AI infrastructure performance across three core dimensions: 98% of the theoretical line rate with low jitter-free latency; strong cross-tenant isolation for concurrent workloads; robust, capacity-proportional bisection bandwidth and 7% latency increase for 10% fabric link failures; and rapid reaction to host and fabric link flaps during LLM training workloads.
Bio
Dr. Richard Graham is a Senior Director at NVIDIA's Networking Business unit. His primary focus is on HPC and AI network software and hardware capabilities for current and future HPC and AI technologies. Prior to moving to Mellanox/NVIDIA, Rich spent thirteen years at Los Alamos National Laboratory and Oak Ridge National Laboratory, in computer science technical and administrative roles, with a technical focus on communication libraries and application analysis tools. He is cofounder of the Open MPI collaboration and was chairman of the MPI 3.0 standardization efforts.
Abstract
To Be Determined
Bio
Hemal Shah is a Distinguished Engineer and Systems/Software/Standards architect in the Core Switching Group (CSG) at Broadcom Inc. Hemal is responsible for the definition of system/software architecture and roadmap of Ethernet NIC product lines. Hemal is one of the key architects of end-2-end Ethernet networking solutions for AI, HPC, storage, and cloud infrastructures. Hemal spearheaded the development of stateless offloads, virtualization, SR-IOV, QoS, RoCE, vSwitch offload, management, and security features of Broadcom NICs. Hemal led the architecture definition of several generations of NetXtreme® E-Series/NetXtreme I server product lines and NetXtreme I client product lines. Hemal has defined the system architecture of RDMA hardware/software solutions for more than two decades. Before joining Broadcom in 2005, Hemal worked at Intel Corporation where he led the development of system architecture of communication processors, 10G Ethernet controllers, and TCP/iSCSI/RDMA/security offloads. Hemal is the lead technical representative/contributor from Broadcom Inc. in the Open Compute Project (OCP) and DMTF. Hemal serves as Senior VP of Technology in the DMTF and a project co-lead of OCP Hardware Management project. Hemal has co-authored several OCP specifications, 70+ DMTF specifications, four IETF RFCs, and 10 plus technical conference/journal papers, and 40+ patents. With over 28 years of experience, Hemal holds a Ph. D. (computer engineering) and M.S. (computer science) degrees from Purdue University, M.S.E.E. degree from The University of Arizona, and B.S. (electronics and communication engineering) degree from Gujarat University, India.
Abstract
To Be Determined
Bios
Mihai Gabriel Constantin is a lecturer and researcher at the AI Multimedia Lab, CAMPUS Research Institute, National University of Science and Technology Politechnica Bucharest, Romania. His research interests are focused on multimedia / multimodal data processing with neural networks, deep and ensemble learning, large language models, and training process description and optimization.
Dan Mihailescu is a senior Software Architect at Keysight Technologies based in Bucharest, Romania, with extensive experience in building networking and security verification systems. Dan currently leads teams at Keysight Technologies Romania focused on developing AI infrastructure emulation solutions and drives collaborations with academia and industry partners.
12:15 - 12:30
Group Photo
12:30 - 1:30
Lunch Break, Posters and Demos (Cont'd)
Abstract
The rapid growth of AI systems is placing unprecedented demands on both networking and communication software. This talk presents Microsoft's experience developing Multipath Reliable Connection (MRC) and MSCCL++, two complementary technologies designed for next-generation AI supercomputers. MRC introduces an endpoint-driven transport architecture that improves resilience and load balancing in large-scale Ethernet AI clusters, while MSCCL++ provides a programmable communication framework for developing optimized collective and communication-intensive workloads.
We discuss how these technologies work together to support modern AI training and inference workloads, including Mixture-of-Experts models, large-scale inference systems, and communication-computation overlap. The talk will cover design principles, deployment experiences, and future directions in transport-aware collective communication, highlighting the opportunities and challenges of building communication stacks for AI systems at extreme scale.
Bio
Dr. Jithin Jose is a Partner Software Engineering Manager at Microsoft, where he focuses on communication systems for large-scale AI and HPC platforms. His work includes high-performance networking, transport protocols, collective communication libraries, and the co-design of software and hardware systems for AI supercomputers. He currently leads engineering efforts in areas such as Multipath Reliable Connection (MRC), MSCCL++, AI cluster networking, and large-scale training and inference infrastructure. Prior to joining Microsoft, he held research and engineering roles at Intel and IBM Research. He has published extensively in high-performance computing, networking, and distributed systems, and earned his Ph.D. from The Ohio State University.
Abstract
As LLM serving scales, the limiting factor shifts from raw compute to how efficiently work — and data — moves across heterogeneous hardware. This talk uses performance modeling to motivate disaggregated inference, where different stages of the model run on the hardware best suited to them and communicate over a shared interconnect. Starting from an interactivity-versus-throughput analysis, we show how disaggregation shifts the Pareto frontier. We then examine three deployment patterns and the distinct demands each places on the interconnect: prefill–decode disaggregation, which is bandwidth-bound; speculative decoding, which needs balanced bandwidth and latency; and attention–FFN disaggregation (AFD), whose per-layer exchange is latency-critical. We close with the systems techniques we are developing to meet these requirements across GPUs and accelerators, drawing on an ongoing hardware–software collaboration.
Bio
Sangamesh Kodge works on efficient Large-Language Model Inference deployment at d-matrix, spanning performance modeling, disaggregated deployment, and the runtime and kernel implementations for Mixture-of-Experts(MoE) model on the Corsair Platform. He earned his PhD from Purdue University under Prof. Kaushik Roy in Nanoelectronics Research Lab, where his research covered in-memory computing, hardware-software co-design, machine unlearning and decentralized learning with publication at venues including TMLR and AAAI. He holds a B. Tech. in Electrical Engineering from Indian Institute of Technology Kharagpur.
Abstract
The TAU Performance System(R) is a versatile performance evaluation toolkit supporting both profiling and tracing modes of measurement. It supports performance evaluation of applications running on CPUs and GPUs and supports runtime-preloading of a Dynamic Shared Object (DSO) that allows users to measure the performance without modifying the source code or binary. This talk will describe how TAU may be used with MVAPICH and support advanced performance introspection capabilities at the runtime layer. TAU's support for tracking the idle time spent in implicit barriers within collective operations in MPI will be demonstrated. TAU also supports event-based sampling at the function, file, and statement level. TAU's support for runtime systems such as CUDA (for NVIDIA GPUs), Level Zero (for Intel oneAPI DPC++/SYCL), ROCm (for AMD GPUs), OpenMP with support for OMPT and Target Offload directives, Kokkos, and MPI allow instrumentation at the runtime system layer while using sampling to evaluate statement-level performance data. Recent advances include support for PC sampling on AMD, Intel, and NVIDIA GPUs and access to hardware performance counters on GPUs. TAU's support for the key MVAPICH features including its support for the MPI Tools (MPI_T) interface with support for setting MPI_T control variables on a per MPI communicator basis. It will also describe TAU's support for MPI's performance and control variables exported by MVAPICH, and its support for instrumentation of OpenMP runtime, and APIs for instrumentation of Python and Julia programs. TAU uses these interfaces on unmodified binaries without the need for recompilation. This talk will describe these new instrumentation techniques to simplify the usage of performance tools including support for an LLVM plugin for selective instrumentation for compiler-based instrumentation, timing synchronization costs in collective operations, rewriting binary files using DyninstAPI, preloading shared objects. The talk will also highlight TAU's analysis tools including its 3D Profile browser, ParaProf, generating native traces for Perfetto.dev, cross-experiment analysis tool, PerfExplorer and its usage with MVAPICH MPI under Amazon AWS , Google Cloud, Azure, OCI, and IBM Cloud using the ParaTools Pro for E4S(TM) image.
Bio
Sameer Shende serves as a Research Professor and the Director of the Performance Research Laboratory at the University of Oregon and the President and Director of ParaTools, Inc. He serves as the technical lead of the E4S, ParaTools Pro for E4S(TM), TAU Performance System(R), and Program Database Toolkit (PDT) projects. His research interests include scientific software stacks, performance instrumentation, compiler optimizations, measurement, and analysis tools for HPC. He received his B.Tech. in Electrical Engineering from IIT Bombay in 1991, and his M.S. and Ph.D. in Computer and Information Science from the University of Oregon in 1996 and 2001 respectively.
3:00 - 3:30
Coffee Break, Posters and Demos (Cont'd)
Abstract
Cornelis 5K/6K fabrics are evolving beyond a single upper-layer protocol model toward a unified transport services layer that can support IPoIB, RoCE, UE, verbs-based applications, and MPI middleware. This talk introduces the Bulk Transfer Service (BTS)/HFISVC architecture as a common services layer for moving data across multiple protocol personalities while preserving the performance expectations of HPC users. Using MPI communication as the primary case study, we will compare conventional verbs/RDMA paths with Cornelis BTS/HFISVC service-layer paths in terms of latency, bandwidth, CPU overhead, and operational tradeoffs.
Bio
Brian Hwang is a kernel software engineer at Cornelis Networks, where he works on low-level networking software for high-performance computing fabrics. His current work focuses on Linux kernel driver development for HFI-based systems, including service-layer infrastructure and protocol support for IPoIB and other fabric-based communication workloads.
Brian received his M.S. degree in Computer Science from Sungkyunkwan University, where he researched high-performance networking, resource virtualization, and storage which led to a paper acceptance at IEEE CLOUD.
Abstract
To Be Determined
Bio
Mr. Rakesh Kumar Yadav has been associated with the Centre for Development of Advanced Computing (C-DAC) since 2006. He is currently serving as a Scientist E in the HPC Technologies Group at C-DAC, where he leads several key initiatives in system software development, particularly for the Trinetra network. His work also involves the adaptation of MPI libraries for Trinetra. His primary areas of interest include High Performance Computing, Parallel File Systems, Programming models, and Performance tuning. Mr. Yadav holds a Bachelor of Engineering in Information Technology from Maharana Pratap College of Technology, Gwalior, India.
4:30 - 4:50
Student Short Talks (OSU)
4:50 - 5:00
Open MIC Session
6:30 - 9:00
Banquet Dinner
Wednesday, August 19, 2026
Abstract
Supercomputing is evolving toward HPC-AI converged scientific computing, moving from peak-performance-oriented systems to integrated platforms for high-precision simulation, mixed-precision AI, data-intensive workflows, and large-scale scientific applications. This talk introduces LineShine as a next-generation converged supercomputing system designed to support the efficient execution of large-scale HPC-AI workloads. The discussion begins with the architectural requirements of convergence, including balanced compute capability, heterogeneous memory, scalable storage, energy efficiency, and application-driven co-design. It then focuses on communication as a decisive factor at extreme scale. As applications scale to millions of cores and integrate numerical simulation with AI models, sustained performance depends on low-latency interconnects, high bisection bandwidth, topology-aware process placement, optimized MPI communication, scalable collective operations, global reduction optimization, and communication-computation overlap. Software is discussed as the system layer that connects architecture, communication, and applications, enabling complex workflows to run efficiently and reliably on production systems. Representative applications demonstrate how architecture and communication co-design translate system capability into real scientific and engineering productivity.
Bio
Dr. Yutong Lu is a Professor in the Department of Computer Science and Engineering at Sun Yat-sen University, China. She also serves as Director of the National Supercomputing Centers in Guangzhou and Shenzhen. Professor Lu specializes in high-performance computing, with research interests spanning advanced computer architecture, programming models, and parallel computing environments. She was the Deputy Chief Designer of the Tianhe-2 supercomputer. She currently serves as Chief Designer of the LineShine exascale supercomputer. Throughout her career, Professor Lu has been dedicated to bridging high-performance computing with major scientific and engineering applications. she has helped make advanced computing capabilities more accessible to a broader community of researchers and industry users. Professor Lu has been recognized as an ISC Fellow and a CCF Fellow. She has also led multiple major research projects supported by the MOST and the NSFC. Her current research focuses on cutting-edge computer architectures and the convergence of advanced AI and HPC systems and applications.
Abstract
Agentic development has progressed rapidly over the past two years, to the point where it makes sense to ask whether AI can write complex communication software by itself. In this talk, I will describe my experience building an MPI library from scratch using AI. With expert guidance on design and testing, AI was able to implement all of MPI-5 from scratch in less than a month, with support for shared-memory, sockets, OFI/libfabric and UCX, with optimized algorithms for message matching, collectives, etc. I will also talk about agentic development of GPU communication software based on NCCL and NVSHMEM, demonstrating that AI is not limited to CPU environments.
Bio
Jeff Hammond is a Distinguished Engineer at NVIDIA in the data center software organization, focused on GPU communications (NCCL and NVSHMEM). He has extensive experience with the design and use of parallel programming models and scientific applications. Jeff’s most notable achievements include the MPI-5 Application Binary Interface standard, development of the MPI-3 one-sided communication software ecosystem, and contributions to the NWChem quantum chemistry project. He received a PhD in Chemistry from the University of Chicago in 2009.
Abstract
Climate, weather, and environmental research often involve huge spatial datasets spread across different institutions, sensor networks, and computing centers. Bringing all this data together in one place is usually not practical because of high communication costs, data-use rules, privacy concerns, and the growing volume of real-time data. In this talk, I will introduce scalable federated and distributed methods for spatial statistical modeling and environmental prediction. With this framework, each site keeps its own data but works together by sharing small statistical summaries or model updates. The approach uses distributed Gaussian process approximations, decentralized optimization, asynchronous aggregation, and fast local computing to handle large-scale analysis and prediction. This method works well for climate modeling, weather forecasting, air-quality mapping, and environmental monitoring, where data comes from many different sensors and organizations. I will also cover challenges such as data differences, efficient communication, handling delayed updates, keeping models consistent, and measuring uncertainty. Real-world examples will show how federated and distributed spatial analysis can deliver scalable, privacy-friendly, and efficient environmental insights without moving raw data.
Bio
Sameh Abdulah received his M.S. and Ph.D. degrees from The Ohio State University, Columbus, USA, in 2014 and 2016, respectively. He is currently a Senior Research Scientist at the Extreme Computing Research Center (ECRC) at King Abdullah University of Science and Technology (KAUST), Saudi Arabia. His research interests span high-performance computing (HPC) applications, big data analytics, climate and weather modeling, large-scale spatial datasets, parallel spatial statistics, algorithm-based fault tolerance, and machine learning and data mining algorithms. Sameh was a member of the KAUST team nominated for the ACM Gordon Bell Prize in 2022 and awarded the prize in 2024 (Climate Track) for their contributions to large-scale climate and weather modeling and prediction.
10:30 - 11:00
Coffee Break, Posters and Demos (Cont'd)
Abstract
The Fast Fourier Transform (FFT) is a fundamental algorithm used in a wide range of applications, including signal processing, molecular dynamics, particle simulations, and many others. As a result, a high-performance three-dimensional (3D) FFT is essential for these applications, particularly on GPU-accelerated systems.
In this talk, we demonstrate how heFFTe, an open-source library for distributed 3D FFTs on GPUs, leverages the communication optimizations provided by MVAPICH to deliver exceptional performance at scale. Because the FFT is inherently a memory-bound computation, communication over the network often dominates the execution time, making heFFTe a realistic and demanding benchmark for MPI libraries.
We provide an overview of heFFTe's capabilities and present performance results obtained with MVAPICH-Plus on large-scale systems equipped with both NVIDIA and AMD GPUs, highlighting the effectiveness of the communication optimizations across multiple GPU architectures.
Bios
Natalie Beams is a Research Assistant Professor in the Innovative Computing Laboratory at the University of Tennessee, Knoxville. Her interests include numerical methods for scientific computing and high-performance computing for scientific applications. Prior to ICL, she was a postdoctoral research associate in computational and applied mathematics at Rice University. She holds a Ph.D. in Theoretical & Applied Mechanics from the University of Illinois at Urbana-Champaign.
Ahmad Abdelfattah, research assistant professor at the Innovative Computing Laboratory at the University of Tennessee, received his PhD in computer science from King Abdullah University of Science and Technology (KAUST) in 2015, where he was a member of the Extreme Computing Research Center (ECRC). His research interests span high performance computing, parallel numerical algorithms, and general purpose GPU computing. He currently serves as the principal investigator of the MAGMA library. Abdelfattah has been acknowledged by NVIDIA and AMD for contributing to their numerical BLAS libraries, cuBLAS and rocBLAS, respectively.
Abstract
This invited talk presents our recent research on collective communication software for next-generation AI and HPC systems. The first part of the talk introduces the development of vendor-neutral collective communication software for heterogeneous AI accelerators, including GPUs, NPUs, and DPUs. In particular, we present a collective communication software stack being developed at ETRI under a Korean government-funded project to support emerging AI accelerators. The proposed software stack provides a unified collective communication abstraction across AI accelerators from different vendors. Through experimental evaluation, we show that the proposed abstraction introduces only negligible overhead while enabling efficient collective communication on heterogeneous accelerator platforms. We further demonstrate its potential to improve data-processing efficiency and application-level performance for LLM inference workloads.The second part of the talk introduces our research on CXL-based collective communication optimization, conducted in collaboration with Prof. D. K. Panda’s group at The Ohio State University. This work investigates how intelligent CXL switches can be used to accelerate MPI collective operations and improve communication efficiency in rack-scale multi-node systems.
Bio
HooYoung Ahn received the Ph.D. degree in the School of Computing from the Korea Advanced Institute of Science and Technology (KAIST), Daejeon, Republic of Korea, in 2016. She is currently a Principal Researcher with the Supercomputing System Research Section, Electronics and Telecommunications Research Institute, Daejeon, Republic of Korea. Her research interests include distributed and parallel computing, artificial intelligence, and high performance computing.
Abstract
To Be Determined
Bio
Mahidhar Tatineni received his M.S. & Ph.D. in Aerospace Engineering from UCLA. He currently leads the User Services group at SDSC as a Computational and Data Science Research Specialist Manager. He has led the support of high-performance computing and data applications software on several NSF and UC funded HPC and AI supercomputers including Voyager, Expanse, Comet, and Gordon at SDSC. He has worked on many NSF funded optimization and parallelization research projects such as MPI performance tuning frameworks, hybrid programming models, big data middleware, and application performance evaluation using next generation communication mechanisms for emerging HPC systems. He has also led tutorials on AI, HPC, and Kubernetes topics at several PEARC and SC conferences. He is co-PI on the NSF funded Expanse HPC system and the Prototype National Research Platform (PNRP) projects at SDSC. He is the PI on a NSF funded category II system Cosmos that will feature AMD Instinct MI300A accelerated processing units (APUs) that feature both CPU and GPU capabilities with a unified memory architecture.
12:30 - 1:15
Lunch Break, Posters and Demos (Cont'd)
Abstract
To Be Determined
Bio
Abstract
To Be Determined
Bios
Soham Ghosh is a Principal HPC Engineer at X-ScaleSolutions. He works on co-designing mvapich with distributed HPC applications, and on integrating parallel AI workflows with robust parameter optimizations. Before joining X-ScaleSolutions, he was an HPC engineer at the National Energy Research Scientific Computing Center at Berkeley Lab, where he built and optimized scientific codes for large scale hybrid systems.
Kyle Schaefer is a Senior Software Engineer at X-ScaleSolutions. Kyle leads the development of a suite of HPC software products including X-ScaleAI, X-ScaleHPC and MVAPICH2-DPU. His expertise lies in conceptualizing, designing, and delivering end-to-end solutions for HPC and AI.
Abstract
To Be Determined
Bio
A veteran of High-Performance Computing (HPC), Dr. Chaudhary has been actively participating in the science, business, government, and technology innovation frontiers of HPC for almost three decades. His contributions range from heading research laboratories and holding executive management positions, to starting new technology ventures. Most recently, he was a Program Director at the National Science Foundation where he was involved in many national initiatives and the Empire Innovation Professor of Computer Science and Engineering at SUNY Buffalo. He cofounded Scalable Informatics, a leading provider of pragmatic, high performance software-defined storage and compute solutions to a wide range of markets, from financial and scientific computing to research and big data analytics. From 2010 to 2013, Dr. Chaudhary was the Chief Executive Officer of Computational Research Laboratories (CRL), a wholly owned Tata Sons company, where he grew the company globally to be an HPC cloud and solutions leader before selling it to Tata Consulting Services. Prior to this, as Senior Director of Advanced Development at Cradle Technologies, Inc., he was responsible for advanced programming tools for multi-processor chips. He was also the Chief Architect at Corio Inc., which had a successful IPO in July, 2000. Dr. Chaudhary was awarded the prestigious President of India Gold Medal in 1986 for securing the first rank amongst graduating students at the Indian Institute of Technology (IIT). He received the B.Tech. (Hons.) degree in Computer Science and Engineering from the Indian Institute of Technology, Kharagpur, in 1986 and a Ph.D. degree from The University of Texas at Austin in 1992.
2:45 - 3:15
Coffee Break, Posters and Demos (Cont'd)
Abstract
To Be Determined
Bio
Dr. Hari Subramoni is an assistant professor in the Department of Computer Science and Engineering at the Ohio State University, USA. His current research interests include high performance interconnects and protocols, parallel computer architecture, network-based computing, exascale computing, network topology aware computing, QoS, power-aware LAN-WAN communication, fault tolerance, virtualization, big data, deep learning and cloud computing. He has published over 100 papers in international journals and conferences related to these research areas. He has been actively involved in various professional activities in academic journals and conferences. Dr. Subramoni is doing research on the design and development of MVAPICH2 (High Performance MPI over InfiniBand, iWARP and RoCE) and MVAPICH2-X (Hybrid MPI and PGAS (OpenSHMEM, UPC and CAF)) software packages. He is a member of IEEE.