Capability Radar

See what a model can do. At a glance.

The Capability Radar maps LLM benchmark performance across all 11 task dimensions in a single visual profile. The shape shows where a model is strong, where it trails, and how two models compare — with a source-attributed benchmark table beneath every chart.

Reasoning Code Gen QA Factual QA Analytical Classification Extraction Structured Routing Creative Writing Rewriting Technical Writing Schema Mapping
Primary modelComparison model
Illustrative comparison across the 11 task dimensions. Actual profiles are generated from benchmark data in the app.

What is a capability radar?

A capability radar — also called a radar chart or spider chart — is a data visualization that displays multivariate data across multiple axes, each representing a different dimension of performance. Each axis radiates from a central point, and values are plotted along each axis and connected to form a shape. That shape is the profile.

Radar charts are used wherever multidimensional comparison matters: athletic performance, product evaluation, risk assessment, and — increasingly — AI model selection. The human visual system recognizes patterns faster than it compares numbers.

Model Optimizer applies this format to LLM benchmark performance across 11 task dimensions — turning benchmark data into a visual capability profile for every evaluated model.

The Capability Radar as decision intelligence

Benchmark tables tell you how a model scored. The Capability Radar shows you what that means.

Every model evaluated in Model Optimizer generates a visual benchmark profile across eleven task dimensions. That profile creates a capability shape — a fast, readable view of where each model is strong, where it trails, and how its performance varies across different kinds of AI work.

The shape is the insight. And when two models are displayed on the same radar, the shape becomes a decision.

See what a model can do. At a glance.

The Capability Radar is more than a summary. It is a visual decision tool — especially when comparing two models.

A model that leads on analytical and factual task types but trails on creative or language-heavy dimensions produces a different shape than a model optimized for rewriting, content generation, or structured outputs. A narrow indentation on two or three dimensions can reveal more about operational suitability than a ranked list of scores alone.

Eleven dimensions. One profile.

The Capability Radar evaluates each model across the 11 task types in the Model Optimizer framework:

  • Reasoning
  • Code Generation
  • QA Factual
  • QA Analytical
  • Classification
  • Extraction
  • Structured Routing
  • Creative Writing
  • Rewriting
  • Technical Writing
  • Schema Mapping

Each dimension is scaled to its benchmark score and plotted as a point on the radar. The resulting shape is the model's capability profile.

Side-by-side model comparison

The Capability Radar supports direct two-model comparison. A primary model is displayed as one profile. A comparison model is overlaid on the same radar, making overlap and divergence immediately visible.

Benchmark tables require comparison. Radar profiles reveal it.

Dimensions where one model leads, dimensions where both models are closely matched, and dimensions where the gap is significant are visible at a glance. That makes cross-model capability evaluation faster, clearer, and more directly useful for model selection decisions.

Benchmark details and source attribution

Every Capability Radar is paired with a Benchmark Details table providing the source context behind the visual profile — task type, benchmark used, score, rank, percentile, and source attribution.

Every score on the radar has a traceable source in the table beneath it — a published benchmark or a Model Optimizer proprietary evaluation developed for task types where no widely adopted benchmark exists.

Task type Benchmark Score Rank Percentile Source
Reasoning GPQA Diamond 78.2 #2 96th GPQA
Code Generation SWE-bench Verified 65.4 #3 91st SWE-bench
QA Factual SimpleQA 88.1 #1 98th OpenAI
QA Analytical MMLU-Pro 84.7 #2 94th TIGER-Lab
Extraction Model Optimizer Eval 90.6 #4 89th Model Optimizer
Schema Mapping Model Optimizer Eval 79.3 #5 82nd Model Optimizer

Illustrative example rows drawn from the 11 task types. Each live radar displays the full 11-dimension table with the source for every score.

For the full evaluation framework, see LLM Benchmarking Methodology.

From visual to decision

The Capability Radar is the visual entry point into Model Optimizer's capability intelligence layer.

  • LLM Benchmarking Methodology — the evaluation framework and governance principles behind every radar score
  • Model Fitness — how radar performance maps to your actual prompt patterns, workload, and cost constraints
  • AI Cost Optimization — using capability profiles to identify lower-cost models that meet performance requirements
  • Prompt Optimization — aligning prompt design to the models whose capability profiles best match each task type

The radar shows you what a model can do. Model Fitness shows you whether it is the right model for the work you actually do.

Those are the two questions every model selection decision requires.

Every model, profiled across 11 dimensions.

Model Optimizer evaluates 36+ models across the same 11 task types — every score traceable to a published benchmark or a Model Optimizer evaluation. Start free and see the capability profile of every model you run.

Get Started Free

Frequently Asked Questions

What is the Capability Radar?

The Capability Radar is a visual benchmark tool that maps LLM performance across 11 task dimensions — Reasoning, Code Generation, QA Factual, QA Analytical, Classification, Extraction, Structured Routing, Creative Writing, Rewriting, Technical Writing, and Schema Mapping — in a single radar chart profile. The resulting shape makes model strengths, weaknesses, and capability gaps visible at a glance.

Can I compare two models on the Capability Radar?

Yes. The Capability Radar supports direct two-model comparison. A primary model profile is displayed, and a second model is overlaid on the same radar. Overlap, divergence, and performance gaps across all 11 dimensions are immediately visible without parsing benchmark tables.

Where do the benchmark scores come from?

Every score on the Capability Radar traces to a specific benchmark source — either a published benchmark such as GPQA Diamond or MMLU-Pro, or a Model Optimizer proprietary evaluation developed for task types where no widely adopted published benchmark exists. The Benchmark Details table beneath every radar displays the source for each dimension.

What is source attribution?

Source attribution means every score displayed in Model Optimizer identifies its origin — the benchmark name, the organization responsible for it, and the evaluation year. Proprietary evaluations are clearly labeled as Model Optimizer evaluations. No score is presented without a traceable source.

How many models are evaluated?

Model Optimizer currently evaluates 36+ models across 11 task dimensions, and the field grows as new models are released. Each model is scored using the most credible available source for each task — a published benchmark where one exists, or a Model Optimizer proprietary evaluation where one does not.

What does the shape of the radar tell me?

The shape of a model's radar profile shows where it is strong, where it trails, and how its capabilities are distributed across task types. A model optimized for reasoning and code generation produces a different shape than one optimized for creative writing and rewriting. When two models are compared on the same radar, the overlap and divergence between their shapes makes the capability differences between them immediately visible — without requiring row-by-row benchmark comparison.