Pre-release reasoning in open-weight LLM shows persistent wrong answer commitment despite task constraints
A minimal probe task shows Qwen3-8B overwhelmingly recommends walking when only driving satisfies the premise ('car must be at the car wash').
A minimal probe task shows Qwen3-8B overwhelmingly recommends walking when only driving satisfies the premise ('car must be at the car wash').
RIMS is a three-stage preference optimization framework designed for small-scale language models (SLMs) in retrieval-augmented generation (RAG) settings.
Open-source MATLAB framework and FAIR-compliant dataset of 1,326 labeled touch gesture sequences from 25 participants released for affective touch recognition research.
A new arXiv pre-print proposes that annotator stress or distress can systematically shift pairwise preference labels in RLHF, introducing a structured confound.
Six LLMs were tested across spatial navigation, clinical triage, and financial allocation tasks to assess risk attitudes.
A new interpretability technique called the Jacobian lens identifies verbalizable representations in LLMs, termed J-space, which exhibit functional properties of a global workspace.
Proposes a unified text-serialization approach to handle multimodal clinical data (free-text narratives, vital signs, lab values) without task-specific fusion architectures.
VarRate introduces a training-free KV cache compression method for long-context LLMs that allocates variable low-rank budgets to tokens based on query salience, avoiding irreversible evictions.
Cura 1T is a healthcare-specialized LLM introduced in an arXiv preprint (arXiv:2607.15314).
GraphDx introduces a multi-agent framework with Perception, Reasoning, and Decision agents to balance diagnostic accuracy and resource costs.
Causal-Audit is a new framework for explicit, auditable causal reasoning in large language models (LLMs) designed for context-free intervention-based question answering.
Anthropic says it has identified a previously undetected internal space in its Claude models that influences reasoning but does not appear in outputs.
Apple ML Research published a paper introducing interactive proof systems for verifying general distribution properties with bounded-depth circuits.
Apple ML Research published a paper introducing doubly sub-linear interactive proofs of proximity (dsIPPs).
A new Apple Machine Learning Research paper distinguishes testing and verification complexity for location-invariant properties of functions.
Apple’s ML Research team proposes an unlearning framework that reduces computational costs by up to 50% by focusing on low-influence data points.
Apple’s ML Research team introduces Visual Concept Inference from Sets (VICIS), a new task to evaluate vision-language models’ (VLMs) ability to infer shared visual concepts from small sets of example images.
Proposes a three-level hierarchical learning architecture for autonomous UAV swarms in search and rescue, integrating reflex-level neuroplasticity, skill-level MARL with GNNs, and strategy-level meta learning with BDI reasoning.
HG-RAG introduces a graph-traversal pipeline that anchors queries to named entities and expands context upward, laterally, and downward through a hierarchical knowledge graph.
IMEX is a new explainability framework for black-box predictive models that quantifies both individual feature contributions and higher-order interactions.
Google DeepMind and Isomorphic Labs announced a joint bioresilience program to prevent model misuse, improve outbreak detection, and accelerate drug discovery.
SPINE, a multi-agent framework for debugging and deploying bimanual robots, improved operationalization success from 75% to 100% in novice vs. baseline comparisons on DOBOT X-Trainer robots.
Researchers introduced OriginBlame (ob), a system for record- and token-level data provenance in AI training datasets.
A new black-box audit method tests whether LLM chain-of-thought steps depend on stated premises by substituting predicates and re-running the model.
A new benchmark, EgoBabyVLM, tests whether vision-language models can learn like babies using headcam footage from infants.
A proof-of-mechanism study fine-tunes Qwen3.6-27B locally to adapt to a financial ontology, achieving a 0.90 grounded rate on 40 held-out Vietnamese financial tasks, matching a GPT-5 frontier baseline.
A new arXiv survey formalizes in-context reinforcement learning (ICRL) under non-stationarity, where environments change and prior context can become stale or misleading.
A new arXiv preprint formalizes optimal market making in zero-fee perpetual futures as a stochastic optimal control problem on a filtered probability space.
Google DeepMind and India’s Atal Innovation Mission launched ATL Saathi, a Gemini-powered web app for educators in robotics labs.
Proposes a structured framework to decompose image-based retinal diagnosis using the Toulmin model of argumentation.
Prompt wrappers that differ only in formatting can alter LLM benchmark scores enough to reverse leaderboard rankings.
Microsoft Research describes a formal verification workflow that uses Rust, Lean, and Aeneas to prove correctness of cryptographic algorithms in SymCrypt.
Proposes PRecG, a pipeline that segments legal judgments by rhetorical roles and constructs segment-level knowledge graphs to capture legal entities and relationships.
Researchers developed a RAG-based system to automate the generation of investor briefs using company reports, SEC filings, and macroeconomic data.
A new arXiv paper introduces CogniConsole, an architecture that externalizes inference-time control for LLMs into a structured interface.
HALO introduces a hybrid adaptive latent-refinement method to improve frozen pretrained language models with minimal extra compute.
AgentKGV introduces a two-stage training strategy—turn-level distillation-based SFT and trajectory-level GRPO—to improve accuracy and cost-efficiency in knowledge graph fact verification.
Anthropic researchers developed a tool called the Jacobian lens (J-lens) to uncover a hidden internal state space in Claude Opus 4.6, dubbed 'J-space'.
Apple ML Research published a paper proposing a training-free diagnostic framework to evaluate on-policy distillation signals at per-token resolution.
Apple’s ML Research team proposes a method to generate videos with synchronized audio from text while aligning both modalities to the input conditions.
Apple’s ML Research introduces Self-Reflective Program Search for Long Context (SRLM), a framework that augments programming-based context interaction with uncertainty-aware self-reflection.
Apple’s ML Research team introduces Temporal Global Policy Optimization (TGPO), an RL algorithm that incentivizes temporal awareness in multimodal video models.
Apple Machine Learning Research describes a new paper accepted at the AI4TCI workshop at ARES 2026 that formalizes behavioral privacy leakage in agentic negotiation systems.
Flint is an open-source visualization language designed to let AI agents produce expressive, visually polished charts from simple, human-editable specifications.
A new arXiv preprint introduces a human-LLM collaborative framework to construct EspanStereo, a Spanish-language stereotype dataset spanning multiple Spanish-speaking countries.
Aurora 1.5 adds 22 new weather variables, hourly temporal resolution, and probabilistic ensemble forecasting to the open Aurora foundation model for Earth-system applications.
A new arXiv preprint proposes reframing AI for formal mathematics as 'research agents' rather than problem-solvers.
Proactive agents could surface relevant, actionable information to workers before they ask, addressing a key limitation of current reactive RAG and agentic systems.
An open-access paper on arXiv introduces an AI-powered tool that links economic (GTAP) and biophysical (APSIM) models to analyze agricultural supply chain disruptions.
A conceptual framework called adversarial social epistemology (ASE) is proposed to analyze how agents distort or under-specify information in densely interactive human–LLM communicative landscapes.
LLMs are integrated into agent-based modeling to enable real-time adaptation to changing conditions.
A new theoretical framework models in-context search as approximate inference over reasoning traces, where self-reflection provides feedback for posterior updates.
A 16-year-old KVM vulnerability (CVE-2026-53359) allows untrusted guest VMs to escape and gain root on host systems.
A multi-agent AI system automates end-to-end bioinformatics manuscript generation with grounded claims and executed experiments.
A novel framework inspired by statistical mechanics models variable dependencies in cyber-physical IoT systems using an undirected energy-based representation.
A new arXiv preprint introduces a multimodal NLP framework designed to detect misinformation and violence-prone dynamics early.
iFLYTEK-Embodied-Omni is a unified multimodal foundation model that jointly models vision, language, and action within a single framework.
Researchers propose FCPA, a training objective to align LLM validators with frequency-corrected generator outputs.
Local pairwise comparisons may not capture how people truly want automated decision rules to behave when they hold multiple, conflicting priorities.
ASK+ addresses the failure of vanilla uncertainty-gated LLM assistance in partially observable reinforcement learning by supplying trajectory-aware context and structured reasoning to small language models.
TopoPrimer is a framework that explicitly incorporates the global topological structure of a series population into forecasting models.
Sparse MoE models route tokens through subsets of experts per layer, but most possible expert paths remain unused despite practical clustering into a small subset aligned with linguistic function.
Apple ML Research proposes compact seq2seq models for ASR error correction, trained on real and synthetic ASR errors.
Apple researchers propose amortized MIPS, a regression-based method to predict vector search solutions directly rather than computing them repeatedly.
RL-finetuned vision-language models (VLMs) suffer large robustness drops under simple textual perturbations like misleading captions or incorrect chain-of-thought traces.
Apple’s Machine Learning Research team introduces MemoryLLM, a method to decouple feed-forward modules (FFNs) from self-attention in transformers.
Apple’s ML Research team introduces VideoFlexTok, a video tokenizer that outputs variable-length, coarse-to-fine token sequences instead of fixed 3D grids.
Microsoft Research introduces Memora, a scalable memory system for AI agents that separates stored content from retrieval methods to balance abstraction and specificity.
Self-organizing multi-agent LLM teams underperform their strongest individual member by up to 41.1% on ML benchmarks.
TokenScope is a new interactive tool for decoder-based LLMs that exposes token-level metrics, attention patterns, and structural information during generation.
Reasoning LLMs generate long chain-of-thought sequences that accumulate large KV caches, increasing decoding latency and limiting throughput.
Google DeepMind and A24 announced a first-of-its-kind research partnership to develop new workflows and techniques for filmmakers.
PACE introduces a modular neuro-symbolic framework that separates neural prediction from symbolic reasoning to generate counterfactual explanations constrained by domain knowledge.
AFR is a constrained coding-agent workflow that proposes and implements candidate federated learning (FL) algorithmic changes, including server aggregation rules and client update schedules.
Wiola is a new small language model architecture built from first principles, sharing no structural lineage with existing model families such as GPT, LLaMA, Mistral, or Falcon.
Loom is a new assisted-writing framework designed to resolve a persistent failure mode in LLM creative writing assistance, where models oscillate between surface-level polishing and uncontrolled plot expansion.
A new arXiv preprint challenges the assumption that persona representations in large language models are invariant across different operational regimes.
A new arXiv preprint proposes "steering vectors" to directly control language model behavior by intervening in latent space.
A new arXiv paper proposes Bounded Morality, a formal framework that models moral cognition as a tradeoff between moral breadth and moral depth under finite computational resources.
A new arXiv paper introduces MMM, a data model designed to address limitations of document-centric knowledge systems by combining normative constraints with free-text labels.
A new paper introduces Constructive Alignment, a paradigm that reframes AI alignment as a control problem over evolving human preference trajectories.
Researchers propose an AI-driven approach to discover reusable simulation models using natural language queries.
A new arXiv preprint introduces a controlled student-teacher protocol to evaluate when natural-language feedback improves agent performance beyond repeated attempts alone.
A new iterative prompt-optimization framework called Contrastive Reflection improves held-out exact-match accuracy for agentic information retrieval from 51.4% to 60.4% on HotpotQA.
A Hugging Face-affiliated team (Dharma AI) argues AI specialization is theoretically inevitable.
A closed-loop framework links evaluation failures to targeted data or training interventions in LLM development.
A new arXiv preprint introduces a theoretical framework for language generation that explicitly tolerates controlled hallucinations.
DiScoFormer jointly estimates density and score from a set of data points in one forward pass without retraining.
A new arXiv preprint introduces an axiomatic evaluation framework to assess latent thought representations in LLMs, independent of downstream benchmark scores.
A position paper on arXiv proposes reserving 'machine unlearning' for dataset-defined deletion where a model’s training influence is removed such that it is approximately indistinguishable from retraining without that data.
A new arXiv pre-print proposes a three-stage training paradigm to enable LLM agents to internalize future-aware planning.
Proposes AI-ModelNet, a framework to interconnect heterogeneous AI models for collaboration and capability sharing.
Personality prompting (e.g., agreeableness) alters multi-agent LLM communication styles but does not uniformly affect task performance.
Researchers introduce generative causal testing (GCT), a method to translate black-box AI models of brain activity into readable explanations.
Talos, an open-source system for automated genomic reanalysis, recovered 90% of in-scope rare disease diagnoses while surfacing only 1.3 candidate variants per patient for expert review.
Helpfulness-oriented post-training significantly degrades mid-trained compassion values in Llama 3.1 8B, while coding-focused training preserves them.
LLMs perform well on text-only statics problems but see accuracy drop when diagrams are introduced.
Eight of ten tested linguistic features produced statistically significant shifts in Llama-3.2-1B's pro-animal-welfare reasoning when used as fine-tuning data.
Researchers propose a data-generation pipeline to isolate linearly scalable features tied to sycophantic behavior in language models.
Refusal in instruction-tuned chat models is gated downstream by a compliant persona direction, not an isolated mechanism.
HierBias introduces a hierarchical, context-conditioned architecture for media bias detection that models inter-sentence dependencies.
Researchers propose a method to automate the generation of challenging benchmark instances for neural relational reasoning using LLMs.
Proposes a full-stack methodology for constructing agentic AI systems, emphasizing layered understanding beyond LLMs alone.
A new arXiv pre-print proposes the Goal-Identity-Configurator (GIC) architecture to clarify the boundary between agentic and agentive systems.
A neuro-symbolic framework called Neuro-Symbolic Drive supervises driving VLAs using rule-grounded reasoning traces extracted from classical rule-based planners.
EXPO-SQL proposes clause-level execution feedback to improve Text-to-SQL generation using large language models.
Annotation budgets for natural language inference should be set by the target metric, not uniformly.
Seven LLMs (4B–671B parameters) fine-tuned on Arabic showed no Semitic-specific transfer in zero-shot reading comprehension across languages.
A weighted ensemble of Google's Gemini 2.5 Pro, Gemma 3 12B, and Gemma 3 27B achieved a 0.74 weighted F1-score and 0.74 accuracy in detecting EQ-5D studies from PubMed abstracts.
Frontier large language models plateau at a 90.8% initial pass rate on the VerilogEval benchmark for hardware design, according to a new arXiv preprint.
Eight diffusion language models were evaluated across eight benchmarks covering reasoning, coding, translation, knowledge, and structured problem solving.
A new arXiv paper proposes a reusable pipeline to measure how completely undergraduate computer science programs cover international curricular guidelines.
Autonomous agentic AI systems introduce new security, privacy, and compliance challenges beyond traditional access control.
Introduces CaVe-VLM-CoT, a modular reflection-based agentic-RAG framework for vision-language models (VLMs) to reduce hallucinations via evidence-grounded reasoning.
NAVI-Orbital is the first system to perform autonomous multi-modal inference onboard a Low Earth Orbit spacecraft using a vision-language model.
Shared-workspace human-AI teams were evaluated across 1,482 sessions using the Collaborative Gym environment and DiscoveryBench tasks.
PromptMN is a domain-specific language that annotates natural language prompts with structured, %-prefixed directives for roles, goals, constraints, and outputs.
MemSlides introduces a hierarchical memory system for personalized presentation agents, separating long-term user profiles, session-level working memory, and tool memory.
RepSelect targets selective representations to achieve deep and robust forgetting in LLMs.
Introduces SkillChain-Gym, a benchmark for reskilling-aware production-inventory control under disruptions.
DivInit is a training-free intervention that improves agentic search by generating diverse first-turn queries from a single model call instead of sampling independent queries.
A new arXiv paper proposes a self-evolving agent that iteratively refines query-rewriting rules to improve legal case retrieval without parameter training.
Gemini 3.5 Live Translate is a new audio model from Google DeepMind that delivers near real-time speech-to-speech translation in over 70 languages.
Google DeepMind released DiffusionGemma, an experimental open model that uses diffusion-based text generation to achieve up to 4x faster inference on dedicated GPUs compared to autoregressive LLMs.
OSCToM, a new approach combining RL and compositional surrogate models, addresses gaps in how LLMs reason about recursive beliefs and conflicting perspectives in complex social settings.
OpenAI announced that one of its models has disproven a central conjecture in discrete geometry related to the unit distance problem.
Researchers evaluated Gemini 3.0 Flash responses to 2,257 patient health queries under three conditions: without PHR context, with basic summaries, and with full clinical notes.
DeepMind's Co-Scientist AI tool analyzed tens of thousands of scientific papers to identify over 20 novel genetic factors that could reverse cellular aging.
Google DeepMind has integrated Street View imagery into Project Genie, its general-purpose world model, allowing it to generate interactive environments anchored to real geographic locations in the United States.
Researchers conducted a systematic evaluation of theory of mind (ToM) improvements in LLMs by introducing an interactive evaluation paradigm that mirrors first-person, dynamic human-AI interactions rather than third-person benchmarks.
OpenAI hosted Parameter Golf, a community competition with 1,000+ participants and 2,000+ submissions focused on AI-assisted research methods
Researchers introduce Auto-Rubric as Reward (ARR), a framework that externalize preference knowledge as prompt-specific quality rubrics before pairwise comparison.
Researchers tested the common assumption that sharp attention maps correlate with trustworthy answers in vision-language models (LLaVA-1.5, PaliGemma, Qwen2-VL) using a unified mechanistic probing pipeline.
Chain-of-thought reasoning in models like DeepSeek-R1 does not eliminate position bias in multiple-choice tasks; instead, longer reasoning chains correlate with stronger bias toward certain answer positions (correlation 0.11–0.41, all p<0.05).
Apple researchers introduce Reward-Variance Policy Optimization (RVPO), a risk-sensitive alignment framework that prevents language models from achieving high scores in easy objectives while failing at critical constraints like safety or formatting.
Microsoft Research constructed geographically grounded, electrically coherent transmission models of the U.S. power grid entirely from open data sources including OpenStreetMap, EIA statistics, and Census data.
Mozilla engineers found 271 Firefox security flaws using Anthropic's Mythos AI model over a two-month period, with 180 classified as sec-high severity and 80 as sec-moderate.
Researchers introduced Annotator Policy Models (APMs), machine learning systems that infer individual annotators' safety policies from their labeling behavior alone, achieving over 80% accuracy.
AlphaEvolve reduced variant detection errors in DNA sequencing by 30% through improvements to DeepConsensus, a Google Research model for correcting sequencing errors, with results validated by PacBio researchers.
Researchers introduced CreativityBench, a benchmark with 14K tasks grounded in a 4K-entity affordance knowledge base containing 150K+ annotations to evaluate creative tool use in LLMs.
Researchers Kumar and Ahuja introduced LOCA, a causal analysis method that identifies minimal sets of interpretable intermediate representation changes needed to restore refusal behavior in jailbroken LLMs.
Apple researchers created a pseudo-annotation pipeline that automatically generates ranked annotations for sign language video, including glosses, fingerspelled words, and classifiers, reducing manual annotation costs at scale.
Apple researchers introduced a two-agent architecture where a specialized reviewer evaluates tool calls before execution, enabling real-time error correction rather than post-hoc fixes.
Apple Machine Learning Research published STARFlow-V, a normalizing flow-based video generation model accepted to CVPR (April 2026).
Researchers from arXiv cs.AI developed a Bayesian framework that calibrates automated evaluation metrics against human judgment to enable confident model replacement decisions with limited manual evaluation data.
Google DeepMind announced a research initiative exploring AI agents that assist patients under clinical supervision, framed as 'triadic care' involving AI, clinicians, and patients.
Researchers presented COSPLAY, a framework where an LLM decision agent retrieves skills from a learnable skill bank while a parallel agent extracts reusable skills from unlabeled rollouts.
Researchers O'Herlihy and Català introduce the Defensibility Index (DI) and Ambiguity Index (AI) to evaluate content moderation systems based on policy consistency rather than agreement with human labels.
Google DeepMind introduced Decoupled DiLoCo, a distributed training architecture that divides compute into asynchronous 'islands' to improve resilience and reduce bandwidth in global data center training
Researchers at Harbin Institute of Technology document a widespread tendency for language models to call external tools unnecessarily, even when they possess adequate internal knowledge to answer questions.