Multi-agent framework lets LLM agents design and run controlled experiments with simulation models
LLM agents can now design and execute controlled experiments using simulation models, not just generate text or code.
LLM agents can now design and execute controlled experiments using simulation models, not just generate text or code.
Apple’s Machine Learning Research team introduces STARFlow2, a unified multimodal generation architecture that combines autoregressive normalizing flows with a frozen pretrained vision-language model stream.
A new dataset of 3,000 cleaned Khmer documents and 300 user-style queries was constructed for semantic search evaluation.
Humans prefer outputs with fewer but more diverse citations, while LLMs show inconsistent citation-related preferences despite lacking access to sources.
KVBoost introduces a chunk-level KV cache reuse system for Hugging Face-compatible decoder models, enabling reuse regardless of where shared content appears in prompts.
A new arXiv paper reviews the phenomenon of 'model collapse' where AI models degrade when trained on synthetic data.
Two update operators—revision-driven update and delayed elaboration—are proposed to explain how AI and human systems incrementally interpret narratives.
Electronic health records (EHRs) often exceed 100,000 tokens, making long-context reasoning difficult for LLMs.
A new arXiv preprint introduces a causal framework to separate representational bias from behavioral bias in language models.
Apple’s ML Research team introduces Internalized Visual Thinking (IVT), a post-training framework for multimodal video reasoning.