Tuesday 6 October · COLM

Team: P1 human fidelity; P2 personas/populations; P3 interactions/emergence; P4 validation. Two sequential poster conversations per person where listed. All times local PDT. Linked titles open the source.
TimeProgramme / countFour-person assignment
08:45–09:00Opening remarksAll four
09:00–10:00Chelsea Finn — Lessons from Robots for Developing Reliable LLMsOptional all: general reliability
10:00–11:00Oral 1; CollabSkill 10:15, intent reasoning 10:30P1/P4 supporting methods; P2/P3 reading
11:00–13:00Poster session 1
140 total; 4 direct
P1: Mind the Sim2Real Gap in User Simulation for Agentic Tasks
Imperial Ballroom #64; then The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues
Imperial Ballroom #72
P2: InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation
Imperial Ballroom #65
P3: Misalignment Contagion: Can a Misaligned Minority Shift Aligned Agents in Multi-Agent LLM Deliberation?
Imperial Ballroom #69
P4: The Illusion of Stochasticity in LLMs
Grand Ballroom #115
13:00–14:30Lunch / notes / follow-upAll four; do not schedule another formal poster block
14:30–15:30Jennifer Pan — Neutrality Without Neutral ModelsAll four: opinion/information systems
15:30–16:30Oral 2 (supporting cognition/fairness)P4 optional; no new direct simulation result assumed
16:30–18:30Poster session 2
144 total; 7 direct
P1: Identity, Cooperation and Framing Effects within Groups of Real and Simulated Humans
Imperial Ballroom #3
P2: Diversifying Multiple Generative Agents by Aligning with Human Populations
Imperial Ballroom #47; then Escaping the Nash Trap: Structural Estimation and Alignment of Strategic Reasoning in Large Language Models
Imperial Ballroom #46
P3: Simulating Organized Group Behavior: New Framework, Benchmark, and Analysis
Imperial Ballroom #4; then Talk is Cheap, Communication is Hard: Dynamic Grounding Failures and Repair in Multi-Agent Negotiation
Franciscan A #76
P4: Small Foundation Models of Human Cognition and Behaviour
Imperial Ballroom #5; then Can AI Truly Represent Your Voice in Deliberations? A Comprehensive Study of Large-Scale Opinion Aggregation with LLMs
Grand Ballroom #111
Scope: 856 unique main papers / six poster blocks; 29 confirmed directly relevant papers. All titles and 203 abstracts screened. Oral/poster repeats are counted once. This short guide is a selection, not the complete reading list.

Main schedule and board locations use saved official extracts; live freshness unconfirmed. Confirm onsite. Additional suggestions are optional, not extra commitments.

Wednesday 7 October · COLM

PDT · Four-person priority routes · Sequential stops within each poster block

TimeProgramme / countFour-person assignment
09:00–10:00John Langford — Industrial Innovation and its implicationsOptional all
10:00–11:00Oral 3; Do Humans and LLMs Diverge in Belief Revision? Evidence from a Bayesian Analysis at 10:45P1/P4: human-comparison method
11:00–13:00Poster session 3
142 total; 4 direct
P1: Do Humans and LLMs Diverge in Belief Revision? Evidence from a Bayesian Analysis
Grand Ballroom #144; then The Story Shapes the Agent: Narrative Priors in LLM Behavior
Imperial Ballroom #51
P2: What Makes a Sale? Simulating End-to-End Seller-Buyer Retail Dynamics with LLM Agents
Franciscan A #78
P3: Commitment To Cooperation With Self-Negotiated Contracts
Franciscan C #96; then SOTOPIA-TOM: Evaluating Privacy and Information Management in Multi-Agent Interaction with Theory of Mind
Franciscan C #91
P4: Compared to What? Baselines and Metrics for Counterfactual Prompting
Imperial Ballroom #41; then CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment
Franciscan B #90
13:00–14:30Lunch / notes / follow-upAll four; do not schedule another formal poster block
14:30–15:30Panel: James Evans, Jeremy Avigad, Rose Yu, Alison GopnikP1/P2 optional; topic not published
15:30–16:30Oral 4 (no confirmed direct paper)Optional / prepare afternoon questions
16:30–18:30Poster session 4
143 total; 6 direct
P1: Multilingual Agent-Based World Modeling for Social Science
Imperial Ballroom #6
P2: Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces
Imperial Ballroom #17
P3: Social World Models
Imperial Ballroom #7
P4: UrbanLLMind: Scalable LLM-Powered Urban Mobility Simulation with Open-Weight Models
Imperial Ballroom #16

Main schedule and board locations use saved official extracts; live freshness unconfirmed. Confirm onsite. Additional suggestions are optional, not extra commitments.

Thursday 8 October · COLM

PDT · Four-person priority routes · Sequential stops within each poster block

TimeProgramme / countFour-person assignment
09:00–10:00Henry Farrell — Large Language Models are Culture MachinesAll four: culture framing
10:00–11:00Oral 5 (no confirmed direct simulation paper)Optional; otherwise reading/follow-ups
11:00–13:00Poster session 5
144 total; 3 direct
P1: Illusory Truth or Mere Exposure? Model-Dependent Repetition Effects in LLM-Based Social Media Simulations
Grand Ballroom #115
P2: Attractor States Emerge in Multi-Turn LLM Conversations
Imperial Ballroom #8
P3: SocialVeil: Probing Social Intelligence of Language Agents under Communication Barriers
Grand Ballroom #140
P4: Illusory Truth or Mere Exposure? Model-Dependent Repetition Effects in LLM-Based Social Media Simulations
Grand Ballroom #115
13:00–14:30Lunch / notes / follow-upAll four; do not schedule another formal poster block
14:30–15:30Jason Eisner — Is NLP System-Building a Machine Learning Problem?Optional; relevance not inferred
15:30–16:30Oral 6; What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks at 16:00P4: construct validity; poster begins at 16:40
16:40–18:30Poster session 6
143 total; 5 direct
P1: High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination
Imperial Ballroom #34; then RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
Imperial Ballroom #40
P2: TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics
Imperial Ballroom #30; then Beyond expert users: agents should help users construct preferences, not just elicit them
Franciscan C #99
P3: MoltNet: Understanding Social Behavior of AI Agents in the Agent-Native MoltBook
Imperial Ballroom #31; then INSIDE the Student’s Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators
Grand Ballroom #103
P4: What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
Grand Ballroom G-6-117; then In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
Imperial Ballroom #39

Main schedule and board locations use saved official extracts; live freshness unconfirmed. Confirm onsite. Additional suggestions are optional, not extra commitments.

Friday 9 October · parallel workshops

PDT · P1/P2: Social Sim · P3/P4: Agent Behavior. Stay in assigned workshop; no transfer required.

Time bandSocial Sim, P1 and P2Agent Behavior, P3 and P4
08:30 to 09:00No scheduled session.08:30 to 09:00. Opening remarks.
09:00 to 10:0009:00 to 09:10. Opening remarks.
09:10 to 09:40. Noah Goodman, Person Simulation

09:40 to 10:15. Spotlights.
09:00 to 09:30. Natasha Jaques, AI Safety is a Multi-Agent Problem

09:30 to 10:00. Sophia Kazinnik, AI Agents as Economic Actors
10:00 to 11:30Spotlights finish at 10:15.
10:15 to 10:45. Zhongyu Wei, Personalized Foundation Model for Social Simulation

10:45 to 11:00. Coffee.
11:00 to 11:30. Slava Jankin, Can Machines Negotiate a Government? LLM Agents, Coalition Formation, and the Validity Problem in Computational Social Science
10:00 to 11:30. Poster Session I
P3: action-grounded evaluation and social dynamics. P4: asset markets and psychometrics. Named routes below.
11:30 to 12:3011:30 to 12:30. Poster session
P1/P2: two primary conversations each.
11:30 to 13:00. Lunch.
12:30 to 13:3012:30 to 13:30. Lunch.Lunch ends at 13:00.
13:00 to 14:00. Spotlight talks
Time bandSocial Sim, P1 and P2Agent Behavior, P3 and P4
13:30 to 14:0013:30 to 14:00. Marwa Abdulhai
Spotlights continue until 14:00.
14:00 to 14:30Logan Cross, Concordia: Generative Agent-Based Modelling for Social Science
Diyi Yang, Delegation, Reliance, and Oversight in Human-Agent Collaboration
14:30 to 15:3014:30 to 15:30. Breakout session
14:30 to 15:00. Break.
15:00 to 15:30. Kawin Ethayarajh, The World Worlds
15:30 to 16:0015:30 to 15:45. Coffee.
15:45 to 17:00. Panel discussion
15:30 to 16:00. Colin Camerer, Modelling Humans: In Silico & In Mice
16:00 to 17:10Panel continues until 17:00.
17:00 to 17:10. Closing remarks.
16:00 to 17:30. Poster Session II
P3: consensus and belief dynamics. P4: empirical validation. Named routes below.
17:10 to 18:00Social Sim has finished.Posters finish at 17:30.
17:30 to 18:00. Closing remarks.

18 workshops investigated: six complete rosters, five partial, seven unavailable. These two are priority choices, not exhaustive coverage. Talk titles/topics marked unpublished remain unknown. Poster/spotlight repeats are not additional papers.

Friday · priority paper conversations

Social Sim: P1/P2, 11:30–12:30. Agent Behavior: P3/P4, I 10:00–11:30; II 16:00–17:30. Two stops per person in each assigned block.

OwnerLinked paper / relevanceWhy attend
P1BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks
direct confirmed
Behavioral-science benchmarks and distributional evaluation
P1A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing
direct confirmed
Human multi-turn belief trajectories and Bayesian latent beliefs
P2Distinguishing Governance Effects from Prompting Artifacts in LLM Pricing Simulations
direct candidate
Economic governance versus prompting confound
P2Everyone Conforms, No One Believes: Pluralistic Ignorance in LLM Agent Populations
direct candidate
Pluralistic ignorance and collective conformity
P3_IEvaluating Generative Agents with Actions Grounded in Socially Distributed Task Environments using Incognita
supporting adjacent
action-grounded social-agent evaluation
P3_IWhat LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates
supporting adjacent
emergent social structure/objectives
P4_IPersona Design and Bubble Formation in LLM-Populated Experimental Asset Markets
direct candidate
persona sensitivity and emergent market behavior
P4_IRethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior
supporting adjacent
Big Five versus behavior-specific TPB; context/persona confounds across 11 models
P3_IIInherited or Formed? The Provenance of Consensus in LLM Agent Societies
direct candidate
emergent versus inherited consensus validity
P3_IIBelief Engine: Configurable and Inspectable Stance Dynamics in Multi-Agent LLM Deliberation
direct candidate
direct inspectable opinion-dynamics tool
P4_IIAugmented Hypothesis Testing with Persona-Based LLM Simulations
direct confirmed
statistically valid use of fallible persona predictions; four real datasets
P4_IICan AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation
direct confirmed
simulated RCT error decomposition and calibration against 67 historical tests
ABSTRACT-confirmed direct work is distinguished from title candidates and adjacent work. Social Sim individual poster attendance is not separately mapped. Ask every author: what human data grounds this model, which held-out test validates it, and does it beat a simpler baseline?