How to use
← OverviewEnter a Hugging Face model ID (e.g. Qwen/Qwen3-8B) and press Analyze.
Auto: uses the checkpoint's native precision (FP8 stays FP8, not expanded to BF16).
Advanced options let you override precision, KV cache dtype, engine, workload preset, and memory utilization.
For gated models (e.g. Llama-3), paste a Hugging Face token in the advanced section.
Results show GPU comparison, memory breakdown, KV cache, topology, concurrency planning, and more.
Confidence labels: EXACT (safetensors verified), HIGH (known architecture), MEDIUM (partial), LOW (analytical estimate only).
Use the tabs to explore different aspects: Overview, Architecture, Memory, KV Cache, Performance, GPU Comparison, Topology, Concurrency, Storage, Assumptions, Formula Audit, AI Explanation.