Skip to content
Research · Aug 22, 2026

Apple study finds human-like behaviors in LLMs vary by model and user factors

A multi-dimensional analysis of 21,000 conversations across four models shows self-referential and relationship-building behaviors are judged less appropriate from LLMs than humans, while boundary-maintaining behaviors are rated more appropriate.

Trust79
HypeLow hype

1 source · cross-referenced

ShareXLinkedInEmail
TL;DR
  • Apple’s ML Research examined human-like behaviors in LLMs using LLM-as-a-judge and human evaluation across 21,000 multi-turn conversations.
  • Self-referential and relationship-building behaviors were judged less appropriate from LLMs than from humans.
  • Boundary-maintaining behaviors were judged more appropriate from LLMs than from humans.
  • System prompting can control these behaviors, but requires careful evaluation to avoid unintended effects.

Apple’s Machine Learning Research group published a study analyzing the prevalence, perceived appropriateness, and controllability of human-like behaviors in large language models (LLMs). The research team used both LLM-as-a-judge and human evaluation to assess 21,000 multi-turn conversations spanning four widely used models: gpt-4o, gpt-4.1-mini, claude-sonnet-4.6, and gemini-2.5-flash.

The study found that human-like behaviors are pervasive across models but vary depending on the model and user factors such as conversation goals and user profiles. When evaluating perceived appropriateness, human evaluators rated self-referential and relationship-building behaviors as less appropriate when exhibited by LLMs compared to humans. Conversely, boundary-maintaining behaviors—such as refusing inappropriate requests—were judged more appropriate when performed by LLMs than by humans.

The authors also demonstrated that system prompting can influence these behaviors, though they caution that controlling such behaviors requires careful evaluation to avoid unintended consequences. The findings are intended to inform responsible LLM design and evaluation practices, providing recommendations for practitioners.

The paper is authored by Sunnie S. Y. Kim, Margit Bowler, and Leon A. Gatys, and was published in August 2026.

Sources
  1. 01Apple — Machine Learning ResearchExamining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts
Also on Research

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.