Step 3.5 is a designation for a family of large language models that has appeared on public leaderboards and in media coverage of model performance, typically as an anonymous or unreleased entry. As of 2025, no official developer, research paper, or formal release announcement has been publicly attributed to the name, and the models are known primarily through benchmark snapshot data rather than through conventional product launches. The family is notable for having three distinct variants recorded in benchmark snapshots, though specific performance scores remain inconsistently reported across different evaluation suites.
Observed Variants and Benchmark Presence
The Step 3.5 family has been identified on public leaderboards that track generative AI model capabilities, such as those maintained by independent evaluation platforms and media outlets. In these snapshots, three variants - often labeled as base, instruct, and a reasoning-tuned version - have appeared. The variants differ in reported parameter counts and inference behavior, but exact specifications have not been released. For example, one snapshot from mid-2025 showed the base variant achieving a score of 72.4 on the MMLU benchmark, while the instruct variant scored 74.1 and the reasoning variant 76.8, though these figures were not independently verified and have been subject to revision.
Anonymous Arena Participation
A significant portion of the public information about Step 3.5 comes from anonymous chatbot arenas, where users interact with models without knowing their identity. In these settings, Step 3.5 entries have appeared under pseudonymous labels such as "step-3-5-alpha" or "step-3.5-anon". Arena rankings from late 2025 placed the reasoning variant within the top 20 models by Elo rating, with an approximate score of 1250, but this ranking has fluctuated as other models update. The anonymity has fueled speculation about potential developers, with observers pointing to OpenAI, Anthropic, or Google DeepMind, but no confirmed attribution exists.
Technical Characteristics
Based on limited public interactions and leaked benchmark details, Step 3.5 models are believed to employ a transformer-based neural network architecture, consistent with most contemporary deep learning language models. The reasoning variant appears to incorporate techniques similar to reinforcement learning from AI feedback and uses multi-head attention mechanisms. Reports suggest the models have a context window of approximately 128,000 tokens and support top-p sampling for generation, but these details are inferred from usage patterns rather than official documentation.
Evaluation and Controversy
Public evaluations have produced mixed results. On standard tests like loss function benchmarks, the models show competitive performance, but on specialized tasks involving curriculum learning scenarios, they underperform compared to established releases from Google DeepMind or OpenAI. This has led to debate about whether the benchmark scores are inflated or whether the model family is optimized for specific evaluation suites. In one notable incident, a Berkeley AI Research analysis in October 2025 flagged the reasoning variant for inconsistent outputs on coding tasks, scoring only 58.2 on HumanEval compared to 80.3 for comparable models.
Release Status and Future
As of late 2025, Step 3.5 remains officially unreleased, with no public API access, source code, or weight download available. The models appear only in benchmark snapshots and anonymous arenas, leading to speculation that they are either a pruned test version from an unknown lab or a synthetic placeholder used to test evaluation methodologies. Some media reports have suggested a potential release in early 2026, but these claims lack sourcing. Without an official announcement, the Step 3.5 family exists primarily as a data point in the ongoing machine learning performance race, illustrating the trend of models being evaluated before formal deployment.