HypeDelta
DigestTopicsClaimsPredictionsReliabilityResearchers
Admin
DigestTopicsClaimsPredictionsReliabilityResearchers

HypeDelta - AI Research Intelligence

Claims Browser

Search and filter through extracted claims from AI researchers.

Search & Filters
All
agents
benchmarks
general
infrastructure
interpretability
multimodal
other
policy
All
critique
fact
hint
opinion
prediction
7d
14d
30d
90d

Showing 61-80 of 94 claims in topic "rlhf"

rlhf
fact
Bearish
academic

RLHF preference tuning drives LLMs to overuse emphatic rhetorical figures like epanorthosis, with the training distribution rich in promotional prose being the main driver rather than the left-to-right generation process

Computation and Language
7/29/2026
Confidence: 80%Source
rlhf
Previous
1235
Page 4 of 5
fact
Neutral
unknown

Removing GRPO's standard deviation normalization transforms catastrophic reward performance from 0% to baseline parity

Machine Learning
7/28/2026
Confidence: 80%Source
rlhf
fact
Bearish
unknown

Potential-based prediction rewards under GRPO drive LLM agents into degenerate absorbing states (the 'dark room' pathology)

Machine Learning
7/28/2026
Confidence: 85%Source
rlhf
fact
Bearish
unknown

Dense per-step supervision with group-normalized RL can destroy policy performance rather than improve it

Machine Learning
7/28/2026
Confidence: 90%Source
rlhf
fact
Bullish
academic

PrefReward outperforms non-personalized and retrieval-based baselines in both generation quality and personalization interpretability on the LongLaMP dataset

Computation and Language
7/28/2026
Confidence: 85%Source
rlhf
critique
Bearish
academic

Most existing personalization approaches for LLMs rely on implicit representations within model parameters, making it difficult to interpret user-specific preferences or handle long-context dependencies

Computation and Language
7/28/2026
Confidence: 75%Source
rlhf
fact
Bullish
unknown

Fusing hidden-state representations from heterogeneous reward models provides a more expressive basis for preference evaluation than scalar-score methods

Computation and Language
7/28/2026
Confidence: 70%Source
rlhf
fact
Neutral
unknown

Existing evaluation approaches fail to adequately capture diverse evaluative criteria underlying human preferences in non-verifiable tasks

Computation and Language
7/28/2026
Confidence: 80%Source
rlhf
fact
Bullish
academic

The algorithm achieves asymptotically optimal regret matching the contextual-bandit lower bound up to logarithmic factors

Machine Learning (Statistics)
7/28/2026
Confidence: 90%Source
rlhf
fact
Bullish
academic

A new algorithm achieves horizon-free regret of O(sqrt(SAK)+S^8A^3) for finite-horizon tabular MDPs, completely removing log H dependence

Machine Learning (Statistics)
7/28/2026
Confidence: 95%Source
rlhf
fact
Bullish
lab researcher

Async OPD with local Monte Carlo next-token sampling achieves ~2x throughput improvements while matching or beating sync OPD accuracy on math tasks

Lewis Tunstall
7/28/2026
Confidence: 80%Source
rlhf
opinion
Bullish
lab researcher

Decoupling generation from learning in a fully async manner is a promising approach for both distillation and GRPO training

Lewis Tunstall
7/28/2026
Confidence: 70%Source
rlhf
fact
Bullish
lab researcher

Monte Carlo sampling with importance corrections remains stable for large staleness of up to 32 steps in async distillation if MC sample size is sufficient

Lewis Tunstall
7/28/2026
Confidence: 70%Source
rlhf
fact
Bullish
academic

Neuron On-Policy Self-Distillation can enable annotation-free post-training by leveraging internal neuron activations for data selection and teacher construction

Machine Learning
7/27/2026
Confidence: 75%Source
rlhf
fact
Bearish
academic

SFT- and GRPO-based self-evolution methods suffer from out-of-domain performance degradation

Machine Learning
7/27/2026
Confidence: 80%Source
rlhf
fact
Bearish
academic

Reward-based on-policy RL methods inflate calibration error in annotation-free self-evolution

Machine Learning
7/27/2026
Confidence: 80%Source
rlhf
fact
Bullish
academic

LF-IBIS enables full Bayesian inference in reinforcement learning settings where environment dynamics lack an explicit or tractable likelihood function

Machine Learning (Statistics)
7/27/2026
Confidence: 85%Source
rlhf
fact
Bullish
academic

Bayesian Reinforcement Learning can address data scarcity challenges by leveraging prior knowledge and sequential belief updates

Machine Learning (Statistics)
7/27/2026
Confidence: 80%Source
rlhf
hint
Bullish
lab researcher

OpenAI is actively hiring post-training experts to improve their Tinker model

John Schulman
7/27/2026
Confidence: 90%Source
rlhf
opinion
Bullish
lab researcher

Fine-tuning with expert judgment data can beat prompting-only approaches by a significant margin, even as general-purpose models improve

John Schulman
7/27/2026
Confidence: 80%Source
Next

Pipeline data may be stale or degraded.

Last synthesis: 2026-09-20. 8,949 pending.