Search and filter through extracted claims from AI researchers.
Showing 41-60 of 451 claims in topic "infrastructure"
Reproducing SmolLM3-3B requires over $700K in training costs
"reproducing SmolLM3-3B needs over \$700K"
"Our best model is trained at a compute cost of less than \$6.9K and approaches Qwen2.5-1.5B performance under our evaluation protocol"
"a cost-efficient, hardware-accessible, and open-source pretraining recipe has long been missing"
"we train a collection of Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision on consumer-grade RTX 5090 GPUs"
"Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities"
"Whole-slide images (WSIs) are central to computational pathology but are prohibitively large, making patch-based processing the practical unit for foundation model inference."
"At scale, however, generating and handling massive numbers of patches on quickly introduces significant I/O and orchestration overhead, often dominating end-to-end performance."
Decoupling I/O, computation, and ingestion enables high-throughput WSI embedding extraction at scale
"We show that decoupling I/O, computation, and ingestion enables high-throughput WSI embedding extraction at scale."
"By characterizing the scaling envelope, we demonstrate that storage dominates beyond moderate concurrency, reframing WSI embedding extraction as a data-centric systems problem rather than a purely compute-bound workload."
"This representation database is compact and reusable for tasks such as retrieval, classification, and few-shot learning, particularly benefiting low-resource environments."
"the all-parallel floor reaches $0.286$ at the final slot on Qwen3-4B, limiting even the best proposal to $71\%$ per-slot acceptance"
"one realised token removes $86$--$100\%$ of this floor, a locality also recovered by an independent mutual-information analysis"
"current drafters remain far above their floors: the final-slot model gap accounts for $43$--$64\%$ of DFlash rejection and $85$--$92\%$ of DSpark's oracle-conditioned rejection"
"Their rejection mixes two losses: missing within-block path information and imperfect modelling of observable information"
About $4.4K is sufficient to reach the performance of Qwen2-1.5B based on the Puro Cost Scaling Law
"the fitted law suggests that about \$4.4K, less than \$5,090, is sufficient to reach the performance of Qwen2-1.5B"
"progress in vision generative AI has been driven by output quality, with hardware evolving reactively to accommodate growing model demands"
"vision generative models are equally needed in applications that operate under strict hardware constraints at the edge, including autonomous vehicles, agricultural sensors, and mobile devices"
"a software-hardware co-design approach, where deployment constraints are considered from the start of the design process, ensuring that the "right model" runs on the "right hardware" to serve the "right application", making generative AI deployment sustainable and accessible across a much broader range of platforms"
LLM compute accounts for 80-90% of query cost in production semantic data processing systems
"In production, LLM compute accounts for $80-90\%$ of query cost"
Each LLM call costs 100,000 to 10,000,000 times more than a relational predicate
"each call costs $10^5-10^7\times$ a relational predicate"
Pipeline data may be stale or degraded.
Last synthesis: 2026-09-20. 8,949 pending.