The FBGEMM Project
The FBGEMM Project is a repository of highly-optimized kernels used across deep learning applications.
The codebase is organized and published as three related packages: FBGEMM, FBGEMM-GPU, and FBGEMM-GenAI. Each package has its own set of features and documentation.
Project Overview
FBGEMM: A low-precision, high-performance matrix multiplication and convolution library for server-side inference. The documentation below provides an overview of FBGEMM, including its features, documentation, and community resources.
FBGEMM_GPU: A collection of PyTorch GPU operator libraries built on top of FBGEMM for training and inference, with focus on recommendation systems applications. Please see
for more information.
FBGEMM_GPU GenAI: A collection of PyTorch GPU operator libraries that are designed for generative AI applications, such as FP8 row-wise quantization and collective communications. Please see
for more information.
FBGEMM (Facebook GEneral Matrix Multiplication) is a low-precision, high-performance matrix-matrix multiplications and convolution library for server-side inference.
The library provides efficient low-precision general matrix multiplication for small batch sizes and support for accuracy-loss minimizing techniques such as row-wise quantization and outlier-aware quantization. FBGEMM also exploits fusion opportunities in order to overcome the unique challenges of matrix multiplication at lower precision with bandwidth-bound operations.
FBGEMM is used as a backend of PyTorch quantized operators for x86 machines:
PyTorch:
https://github.com/pytorch/pytorch/tree/master/aten/src/ATen/native/quantized/cpu
See the full
for more information on building, installing, and developing with FBGEMM, as well as the most up-to-date support matrix and API documentation for this library.
What's New?
New Features and Recent Improvements
(January, 2020)
Citation
For a high-level overview, design philosophy and brief descriptions of various parts of FBGEMM please see
.
For those looking for the appropriate article to cite regarding FBGEMM, we recommend citing our
:
@article{fbgemm, title={FBGEMM: Enabling High-Performance Low-Precision Deep Learning Inference}, author={Khudia, Daya and Huang, Jianyu and Basu, Protonu and Deng, Summer and Liu, Haixin and Park, Jongsoo and Smelyanskiy, Mikhail}, journal={arXiv preprint arXiv:2101.05615}, year={2021} } Join the FBGEMM community
For questions, support, news updates, or feature requests, please feel free to:
File a ticket in
Post a discussion in
Reach out to us on the #fbgemm channel in
For contributions, please see the
file for ways to help out.
Community Ports
Huawei Ascend NPU
provides community-maintained NPU implementations for selected FBGEMM_GPU operators.
Prebuilt wheels: Install from
with python -m pip install fbgemm-ascend, or download a wheel from
.
Build from source: See the
for source build instructions.
License
FBGEMM is BSD licensed, as found in the
file.