A lightweight, general-purpose framework for evaluating GPU kernel and benchmark.
-
Updated
Aug 21, 2026 - Python
A lightweight, general-purpose framework for evaluating GPU kernel and benchmark.
Deep Learning Inference benchmark. Supports OpenVINO™ toolkit, TensorFlow, TensorFlow Lite, ONNX Runtime, OpenCV DNN, MXNet, PyTorch, Apache TVM, ncnn, PaddlePaddle, etc.
First public benchmark of llama.cpp speculative decoding on Qwen3.6-35B-A3B with a single RTX 3090 (post PR #19493 merge, 2026-04-19). 19 configurations covering ngram-cache, ngram-mod, and classic draft with vocab-matched Qwen3.5-0.8B. Finding: no variant achieves net speedup on Ampere + A3B MoE. Raw JSON, plots, full reproducibility.
WER and inference-cost benchmark for Indian-English ASR, and an audit of the evaluation itself: 9 systems x 3 corpora, cluster-bootstrap CIs with Holm correction, a reference-artifact taxonomy confirmed by human re-transcription, and a speaker-disjoint fine-tuning study across 6 seeds x 3 sizes.
Benchmark for GB10 - Nvidia DGX Spark
Benchmark your GPU against any GGUF model and contribute to the public leaderboard. Measures throughput, TTFT, ITL, and VRAM limits across quantizations and context sizes.
Vram budgeting and tail-latency simulation for three models sharing one triton gpu
InfernoBench: open-source AI model installers, local inference benchmarks, and reproducible browser labs. Wan2.1 and Sulphur-2 included.
Local benchmarking tool to explore Vision Transformer scaling and WSI-level inference constraints under real hardware settings.
Local-first Edge AI inference run registry and comparability checker for multi-target benchmark evidence.
Real-time EEG neurological classification benchmark with latency-accuracy tradeoffs, OpenNeuro data, ONNX inference, and parallel FFT speedups.
Benchmark speculative decoding performance for Qwen3.6-35B-A3B on an RTX 3090 GPU using llama.cpp to evaluate model throughput and structural regressions.
Dynamic batching inference server with FastAPI, dynamic padding, and benchmark visualizations
Measuring the KV-transfer tax of disaggregated prefill/decode LLM serving (vLLM + NIXL, 4x A10G): on PCIe-only hardware, disagg loses to plain data-parallel replication — measured, committed, reproducible.
Reproducible benchmarking framework for signal-processing algorithms and ML inference supporting real and synthetic ECG inputs.
Add a description, image, and links to the inference-benchmark topic page so that developers can more easily learn about it.
To associate your repository with the inference-benchmark topic, visit your repo's landing page and select "manage topics."