Skip to content
Change the repository type filter

All

    Repositories list

    • dynamo

      Public
      A Datacenter Scale Distributed Inference Serving Framework
      Rust
      Other
      1.4k7.6k214664Updated Jul 24, 2026Jul 24, 2026
    • aiperf

      Public
      AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solution.
      Python
      Apache License 2.0
      1294633234Updated Jul 24, 2026Jul 24, 2026
    • Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and improve overall performa…
      Python
      Apache License 2.0
      45861623Updated Jul 24, 2026Jul 24, 2026
    • Offline optimization of your disaggregated Dynamo graph
      Python
      Apache License 2.0
      1413724138Updated Jul 24, 2026Jul 24, 2026
    • OpenEngine is a vendor-neutral gRPC protocol for coordinating inference engines and distributed frameworks.
      Apache License 2.0
      1181Updated Jul 23, 2026Jul 23, 2026
    • Rust
      Apache License 2.0
      109011Updated Jul 23, 2026Jul 23, 2026
    • nixl

      Public
      NVIDIA Inference Xfer Library (NIXL)
      C++
      Other
      3751.1k62191Updated Jul 23, 2026Jul 23, 2026
    • grove

      Public
      Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling
      Go
      Apache License 2.0
      782433735Updated Jul 22, 2026Jul 22, 2026
    • Enhancement Proposals and Architecture Decisions
      Apache License 2.0
      1911458Updated Jul 20, 2026Jul 20, 2026
    • TypeScript
      Apache License 2.0
      4502Updated Jul 10, 2026Jul 10, 2026
    • velo

      Public
      Rust
      Apache License 2.0
      2604Updated Jul 7, 2026Jul 7, 2026
    • aitune

      Public
      NVIDIA AITune is an inference toolkit designed for tuning and deploying Deep Learning models with a focus on NVIDIA GPUs.
      Python
      Apache License 2.0
      3127920Updated Jul 6, 2026Jul 6, 2026
    • FlexTensor is a tensor offloading and management library for PyTorch that enables running large models on limited GPU memory by intelligently offloading tensors…
      Python
      Apache License 2.0
      1310900Updated Jun 3, 2026Jun 3, 2026
    • .github

      Public
      3101Updated Aug 21, 2025Aug 21, 2025
    ProTip! When viewing an organization's repositories, you can use the props. filter to filter by custom property.