Copyright Notice:

The documents distributed by this server have been provided by the contributing authors as a means to ensure timely dissemination of scholarly and technical work on a noncommercial basis. Copyright and all rights therein are maintained by the authors or by other copyright holders, notwithstanding that they have offered their works here electronically. It is understood that all persons copying this information will adhere to the terms and constraints invoked by each author's copyright. These works may not be reposted without the explicit permission of the copyright holder.

Publications of SPCL

A. Ivanov, N. Dryden, T. Ben-Nun, T. Schneider, S. Ashkboos, T. Hoefler:

 STen: Productive and Efficient Sparsity in PyTorch

(ACM Trans. Archit. Code Optim.. presented in New York, NY, USA, Association for Computing Machinery, ISSN: 1544-3566, May 2026, Just Accepted )

Publisher Reference

Abstract

As deep learning models grow, sparsity is becoming an increasingly critical component of deep neural networks, enabling improved performance and reduced storage. However, existing frameworks offer poor support for sparsity. Specialized sparsity engines focus exclusively on sparse inference, while general frameworks primarily focus on sparse tensors in classical formats and neglect the broader sparsification pipeline necessary for using sparse models, especially during training. Further, existing frameworks are not easily extensible: adding a new sparse tensor format or operator is challenging and time-consuming. To address this, we propose STen, a sparsity programming model and interface for PyTorch whose key design insight is the decoupling of sparsity layouts, operators, and sparsifiers into composable, first-class abstractions that users can independently define and combine. An automatic dispatch mechanism selects the best available sparse implementation and transparently falls back to dense execution, allowing STen to support virtually all sparsification methods while enabling rapid prototyping without sacrificing performance for optimized paths. We demonstrate the versatility of STen by expressing existing sparsification techniques within its abstraction, achieving a code size reduction of over 2 ×. Finally, we develop a novel, high-performance grouped n:m sparsity layout for CPU inference at moderate sparsity, accelerating end-to-end BERTBASE inference by up to 3.2 ×. STen brings high performance and ease of use, making sparsity readily accessible for existing PyTorch models.

Documents

download article:
 

BibTeX

@article{ivanov2026sten,
  author={Andrei Ivanov and Nikoli Dryden and Tal Ben-Nun and Timo Schneider and Saleh Ashkboos and Torsten Hoefler},
  title={{STen: Productive and Efficient Sparsity in PyTorch}},
  journal={ACM Trans. Archit. Code Optim.},
  year={2026},
  month={05},
  location={New York, NY, USA},
  publisher={Association for Computing Machinery},
  issn={1544-3566},
  note={Just Accepted},
  doi={10.1145/3815424},
}