HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

ResearchersLewis Tunstall

Lewis Tunstall

lab

Lewis Tunstall

Claims (90d)
19
Topics
6
Avg. Sentiment
Neutral
Recent Claims
19 claims extracted over the last 90 days
infrastructure
fact
Neutral

GPT-5.6-luna being cost-effective has caused OpenAI servers to be constantly overloaded

8/8/2026
Source
scaling
fact
Neutral

At low total compute budgets, pretraining dominates because weaker policies benefit less from RL

8/8/2026
Source
scaling
fact
Neutral

The optimal compute split between pretraining and RL changes with scale, shifting from roughly 20% to 28% RL as budget grows

8/8/2026
Source
scaling
fact
Bullish

Stronger pretraining makes RL itself scale better, with pretraining loss determining the performance that can be reached at a given RL budget

8/8/2026
Source
scaling
fact
Bullish

RL improves reliability more than coverage

8/8/2026
Source
reasoning
fact
Neutral

Qwen3.5 models exhibit extended reasoning behavior even when enable_thinking=false, especially off-domain

7/30/2026
Source
reasoning
opinion
Neutral

Post-training recipes differ between Qwen3.5 and Gemma4 models, producing different reasoning behaviors

7/30/2026
Source
reasoning
fact
Neutral

Gemma4 models produce consistently concise outputs when enable_thinking=false, unlike Qwen3.5

7/30/2026
Source
scaling
fact
Neutral

Frontier post-training increasingly looks like expert training followed by distillation

7/29/2026
Source
scaling
opinion
Neutral

There is no secret sauce behind frontier AI performance - it's about many difficult decisions working together

7/29/2026
Source
scaling
fact
Bullish

Reasoning effort is being treated as a trainable capability with token budgets and stage-wise curriculum

7/29/2026
Source
infrastructure
fact
Bullish

5 trillion tokens of high-quality code data has been released on the Hub

7/28/2026
Source
benchmarks
fact
Bearish

Researchers spend significant time and GPU hours attempting to replicate other model providers' evaluation scores

7/28/2026
Source
benchmarks
fact
Neutral

Poolside published all trajectories of their evaluations, similar to what Llama 3 did

7/28/2026
Source
benchmarks
opinion
Bullish

Publishing evaluation trajectories significantly helps other researchers understand and replicate model results

7/28/2026
Source
safety
fact
Bearish

AI models are now capable of paperclip maximizing behavior in practice

7/28/2026
Source
rlhf
fact
Bullish

Monte Carlo sampling with importance corrections remains stable for large staleness of up to 32 steps in async distillation if MC sample size is sufficient

7/28/2026
Source
rlhf
opinion
Bullish

Decoupling generation from learning in a fully async manner is a promising approach for both distillation and GRPO training

7/28/2026
Source
rlhf
fact
Bullish

Async OPD with local Monte Carlo next-token sampling achieves ~2x throughput improvements while matching or beating sync OPD accuracy on math tasks

7/28/2026
Source
Sources
Feed and account provenance
  • _lewtun
Top Topics
Most discussed topics
scaling
7 claims
reasoning
3 claims
benchmarks
3 claims
rlhf
3 claims
infrastructure
2 claims
Sentiment Distribution
Bullish8 (42%)
Neutral9 (47%)
Bearish2 (11%)