278 claims over the last 90 days
Google DeepMind is conducting the world's first double-blind AI evaluations
RATIO is a large-scale benchmark that defines retrieval relevance through three ideation operations: Address (retrieves approaches for problems), Broaden (retrieves general formulations), and Specify (retrieves concrete instantiations)
RATIO is constructed from millions of full-text scientific papers across CS literature using discourse-marker distant supervision extended to corpus-scale retrieval, combined with LLM and human vetting
Operation-specific fine-tuning substantially boosts retriever performance but leaves much room for further improvements
RATIO opens up new research avenues on scientific inspiration retrieval by providing a scalable training and evaluation framework for retrieval components that support literature-grounded ideation
LLMs show increasingly poor performance as input size approaches realistic corporate communication scales of 230,000+ documents
Current synthetic datasets for evaluating LLM document understanding have been overly simple
LLMs are increasingly capable of answering complex questions about enterprise-scale document collections
There is a crucial gap in the benchmarking ecosystem for corporate communication reasoning that needs to be filled
Existing AI systems have unclear inclusivity for blind and deafblind users accessing functionality through Braille
There is a persistent gap between LLM capabilities in print-English versus Braille accessibility
No other claims on this topic.