Why Are Some Models Gray and "Not Rated" in the LMArena Timeline?
Anyone tracking the evolution of large language models (LLMs) on the LMArena text leaderboard will notice a curious feature: some models appear gray and marked as "not rated" in the timeline. If you’ve ever wondered what these visual cues mean, or why certain models seem to hover in an ambiguous zone despite heavy marketing, this post is for you.
Drawing on LMArena’s leaderboard with style control and the comprehensive Hugging Face dataset: lmarena-ai/leaderboard-dataset, we’ll unpack the story behind verified release dates vs announcements, the power of blind-vote preference as a reality check, the rapid shipping cadence across 15+ labs, and why point releases will dominate the 2026 landscape. Along the way, we’ll touch on “gated” releases and one very notable model that never appeared on LMArena at all.
The Gray Zone: What “Not Rated” Really Means
On LMArena’s timeline, models rendered in gray and tagged “not rated” are those that have not yet received a verified or publicly assessable quality score. This usually means one of several things:
- The model’s release date is uncertain or only loosely tied to a marketing announcement rather than a shipped product.
- The model was never benchmarked through LMArena’s blind-vote preference system.
- The model is gated—accessible to only a handful of users or organizations and never broadly evaluated.
- The release is a “point” or minor update without independent re-ranking or consensus scoring.
In other words, gray “not rated” status is a transparency and rigor flag. It warns users not to treat these entries as direct performance rivals to actively scored models but rather as placeholders or indications that more data is needed.
Verified Release Dates vs. Marketing Announcements
One major source of confusion in the LLM space is the difference between marketing announcements and verified shipping dates. Marketing teams often announce models months before they are available, sometimes as teasers or to style-control the narrative. These announcements generate buzz but don’t guarantee any real-world impact or public accessibility.
By contrast, LMArena aims to track the actual shipping or verified availability date. This is when the model can be evaluated blind by crowd raters or integrated into benchmarks. Only with this confirmation does LMArena assign a full rating and feature the model prominently on leaderboards.
This approach, while conservative, protects against hype-driven artifacts. We’ve seen many announcements that never transitioned to shipped products—models that remain gray or “not rated” indefinitely.
Example: Older 2023-2024 Models
Several older 2023-2024 models silently occupy gray spots on the timeline. For instance, some early-stage prototypes by lesser-known labs were announced with fanfare but never released in a way that allowed detailed benchmarking. Others were released privately or under restricted licenses, hence gated.

Thanks to the Hugging Face dataset’s detailed logs, we can cross-reference release dates with benchmark dates to confirm gaps. Many models failed to accumulate enough blind-voter evaluations, causing their perpetual “not rated” status.
Blind-Vote Preference: The Reality Check
One of LMArena’s key innovations is the blind-vote preference system. Instead of relying solely on numeric benchmarks—which can be cherry-picked or skewed—evaluators are shown paired model outputs without labels and asked which they prefer.

This approach minimizes bias from brand hype or preconceptions. It’s a more honest, human-centric way to assess language model quality. Models that have not undergone sufficient blind voting remain “not rated.” This helps prevent premature elevation based on announcements alone.
Why This Matters
- Regressions get caught early: If a point release performs worse, blind votes expose it immediately.
- Model comparisons stay current: The system needs a continuous flow of votes to keep ratings fresh.
- Consumer-facing transparency: Users can see models that aren’t yet trusted with full ratings.
The Faster Shipping Cadence Across 15+ Labs
Compared to earlier years, 2025 and 2026 show a clear acceleration in model rollouts. LMArena now tracks 15+ labs shipping multiple iterations rapidly.
This creates a dense timeline where new versions regularly push scores higher, but it also challenges the rating process. Some models ship as experimental “point” releases https://suprmind.ai/hub/ai-models-index/ or internal updates, meaning they may be implemented but lack the volume or diversity of votes for full rating.
These rapid releases explain the increase of gray “not rated” entries in recent months. The community and LMArena's systems need time to catch up with evaluating every variant as it emerges.
Point Releases Dominating 2026
The trend in 2026 points towards “point releases” rather than massive new model launches. These are refinements—better tuning, data updates, or efficiency improvements rather than architecture overhauls.
From an evaluation perspective, point releases often don't justify a new rating on their own. Instead, they update existing ratings or remain gray until re-evaluated through blind voting cycles.
This means the timeline’s gray zone might grow temporarily even as model quality continues to improve behind the scenes.
Gated Releases: The Hidden Models
Another reason some models never receive ratings or appear gray is gating. Some labs restrict models to select partners, enterprise clients, or invite-only beta users. They effectively prevent broad benchmarking or transparent voting.
LMArena has to respect these access limitations, so even eagerly anticipated “unicorn” models can remain invisible in leaderboards. Access gating is often a business or compliance decision rather than a technical one.
One Model That Never Appeared on LMArena
A unique case is the mysterious model announced by a top-tier lab in mid-2024 that never materialized publicly. Despite high-profile demos and aggressive marketing, this model didn’t ship or open for evaluations. Consequently, it never appeared on LMArena’s timeline or Hugging Face datasets.
This is a high-profile “not rated” anomaly that reminds us to distinguish announcement-driven hype from verified shipping and rating. It also fuels ongoing debates about transparency and responsibility in AI development.
Summary: Why You Should Care
Models marked gray and “not rated” in the LMArena timeline are not failures or necessarily inferior. They are indicators—warnings to look deeper or proceed cautiously before trusting performance claims. These entries flag:
- Models with unverified or purely marketing-driven release dates
- Models without broad community blind-voting evaluations
- Gated or restricted models inaccessible for independent verification
- Point releases that await further assessment or aggregate rating updates
Understanding these nuances helps avoid the common benchmark pitfalls: cherry-picked numbers, “hand-wavy” claims of improvement, and outdated single leaderboard snapshots.
Practical Tips for Tracking Models on LMArena
- Check release dates: Prefer verified shipping dates over announcement dates.
- Look at rating status: Gray and “not rated” means treat with a grain of salt.
- Follow blind-vote trends: These give the most honest insight into quality.
- Watch for gating: If a model is inaccessible, automated benchmarks won’t be reliable.
- Expect more point releases in 2026: Smaller updates won’t always get new ratings immediately.
Final Thoughts
The LLM landscape is fast, noisy, and often blurrier than it seems. The gray “not rated” models in LMArena’s timeline serve as a much-needed reality check amid all the announcements and hype. By grounding yourself in verified data, blind-vote preferences, and thoughtful scrutiny of release cadence, you’ll stay ahead of the curve on real model performance — rather than chasing mirages.
For the sharpest tracking, explore the Hugging Face leaderboard dataset alongside LMArena’s timeline and keep a close eye on the evolving rollout rhythms through 2026 and beyond.