Catch local LLMs spilling into the CPU. One run shows GPU placement, prefill, decode, memory, power and a verdict.
-
Updated
Aug 23, 2026 - Python
Catch local LLMs spilling into the CPU. One run shows GPU placement, prefill, decode, memory, power and a verdict.
This is a uer-friendly Python codebase designed for stress testing of Nvidia GPUs, intel CPUs, and AMD CPUs in various modes
An stress and benchmark utility for NVIDIA GPUs. Measures performance across various precisions (FP64, FP32, TF32, FP16, INT8) and monitors real-time vitals like power, temperature, and clock speeds.
nvProbe — Open-source NVIDIA GPU benchmark suite for CUDA workload automation, Slurm HPC cluster profiling, and MLPerf reporting
PCBench is a versatile Python-based system performance benchmarking tool designed to empower users with insights into their hardware's capabilities. Whether you're a tech enthusiast, a PC gamer, or a developer optimizing your code, PCBench provides comprehensive benchmarking for both CPUs and GPUs.
Benchmark CPU, Benchmark GPU, Storage, RAM using Python
High-performance GPU benchmarking tool built with Vulkan, CUDA, and ImGui — featuring real-time physics simulation, custom rendering, and modular engine architecture.
Re-engineered version of the OpenDwarfs benchmark suite, for compatibility with modern platforms.
Advanced benchmark harness for AI agent inference workloads on AMD GPU and ROCm cloud infrastructure
Reproducible SRAM surrogate simulation benchmark with CPU/CUDA lanes, fidelity validation, and GPU portability architecture.
Local LLM benchmarks on an RTX 5070 Ti 12GB laptop GPU. Speed, coding, reasoning, and tool call accuracy across GGUF models with llama.cpp.
🌌 High-performance WebGL Stress Test. Advanced Raymarching fractal engine with real-time RGB shading and kernel injection.
A code to benchmark GPU performance on different models
**Kernel-V8** is a high-performance GPU benchmarking engine built on the WebGL2 API. By rendering a complex 8th-order **Mandelbulb** fractal in real-time, it generates intense arithmetic workloads to evaluate the stability, thermal throttling, and peak compute throughput of modern graphics hardware.
Reproducible Docker Compose + ComfyUI experiments for MiniMax-H3 video generation.
Pinned llama.cpp benchmark of Qwen3.8-27B on one RTX 3090. Built-in MTP at n-max 2: +59.8% [+57.0, +62.8] server-reported decode on a purposive 25-prompt suite; DFlash2 (PR #27342) +51.9%. Telemetry suggests ~35-37% lower energy, uncalibrated. 23-25 of 25 prompts diverge from serial greedy by token 1600.
AI video generation on RunPod (MiniMax H3 via ComfyUI) assembled with a mechanically-enforced editing grammar
Standalone C++17 SYCL benchmarks for arithmetic, joint-matrix, device-memory, and USM transfer throughput without external compute libraries.
🏆 Which 3DGS renderer is fastest? Which compression is best? We measured them all — on the same GPU, same scenes, same protocol.
GPU vs CPU performance benchmarking for PyTorch and JAX. Works on AMD ROCm, DirectML, CUDA, MPS, CPU. Optimized for RX 5700 XT in WSL2.
Add a description, image, and links to the gpu-benchmark topic page so that developers can more easily learn about it.
To associate your repository with the gpu-benchmark topic, visit your repo's landing page and select "manage topics."