A lightweight, general-purpose framework for evaluating GPU kernel and benchmark.
-
Updated
Aug 8, 2026 - Python
A lightweight, general-purpose framework for evaluating GPU kernel and benchmark.
Deep Learning Inference benchmark. Supports OpenVINO™ toolkit, TensorFlow, TensorFlow Lite, ONNX Runtime, OpenCV DNN, MXNet, PyTorch, Apache TVM, ncnn, PaddlePaddle, etc.
First public benchmark of llama.cpp speculative decoding on Qwen3.6-35B-A3B with a single RTX 3090 (post PR #19493 merge, 2026-04-19). 19 configurations covering ngram-cache, ngram-mod, and classic draft with vocab-matched Qwen3.5-0.8B. Finding: no variant achieves net speedup on Ampere + A3B MoE. Raw JSON, plots, full reproducibility.
WER and inference-cost benchmark for Indian-English ASR, and an audit of the evaluation itself: 9 systems x 3 corpora, cluster-bootstrap CIs with Holm correction, a reference-artifact taxonomy confirmed by human re-transcription, and a speaker-disjoint fine-tuning study across 6 seeds x 3 sizes.
Benchmark for GB10 - Nvidia DGX Spark
Benchmark your GPU against any GGUF model and contribute to the public leaderboard. Measures throughput, TTFT, ITL, and VRAM limits across quantizations and context sizes.
InfernoBench: open-source AI model installers, local inference benchmarks, and reproducible browser labs. Wan2.1 and Sulphur-2 included.
Real-time EEG neurological classification benchmark with latency-accuracy tradeoffs, OpenNeuro data, ONNX inference, and parallel FFT speedups.
Local-first Edge AI inference run registry and comparability checker for multi-target benchmark evidence.
Local benchmarking tool to explore Vision Transformer scaling and WSI-level inference constraints under real hardware settings.
Benchmark speculative decoding performance for Qwen3.6-35B-A3B on an RTX 3090 GPU using llama.cpp to evaluate model throughput and structural regressions.
Reproducible benchmarking framework for signal-processing algorithms and ML inference supporting real and synthetic ECG inputs.
Measuring the KV-transfer tax of disaggregated prefill/decode LLM serving (vLLM + NIXL, 4x A10G): on PCIe-only hardware, disagg loses to plain data-parallel replication — measured, committed, reproducible.
Dynamic batching inference server with FastAPI, dynamic padding, and benchmark visualizations
Add a description, image, and links to the inference-benchmark topic page so that developers can more easily learn about it.
To associate your repository with the inference-benchmark topic, visit your repo's landing page and select "manage topics."