Human-AI Complementarity for Decision Making
2026 Academic Workshop Program Details
Bidirectional Human-AI Alignment: From Static Preferences to Dynamic Human-AI Complementarity
Tiffany Knearem, Aldat Systems Inc.
Hua Shen, NYU Shanghai
Jenny Liang, 91视频
Current alignment paradigms largely optimize AI systems for static human preferences and single-turn evaluations. Yet real deployments involve humans and AI co-adapting over trajectories of interaction: human values shift, model behavior drifts, and complementarity emerges (or fails) longitudinally. This 2-hour, hands-on tutorial introduces Bidirectional Human-AI Alignment, a framework that treats alignment as a dynamic, mutual process: aligning AI with humans (integrating human values and feedback into training and evaluation) and aligning humans with AI (enabling people to explain, audit, and effectively collaborate with AI systems). The tutorial builds on our NeurIPS 2025 tutorial on Human-AI Alignment and the ICLR/CHI BiAlign workshop series.
The tutorial has four parts. (1) Foundations: we synthesize a systematic review of 400+ alignment papers across ML, NLP, and HCI into a unified framework, and clarify what "alignment" means for human-AI complementarity in decision making. (2) Methods: state-of-the-art techniques for value elicitation and specification, preference optimization beyond static RLHF, and evaluation of value-action gaps and multi-turn interaction alignment. (3) Practice: an industry perspective on deploying aligned large models and multimodal interactive systems at scale, including large model quality evaluation and society-centered AI applications. (4) Hands-on: guided Colab exercises in which participants measure the value alignment and interaction alignment of LLMs using open-source tools and benchmarks, and prototype dynamic-alignment evaluations for their own decision-making tasks.
Learning objectives: participants will (a) understand the limits of static, single-turn alignment for achieving human-AI complementarity; (b) acquire a taxonomy of current alignment methods spanning ML and HCI; (c) gain practical experience evaluating the alignment of deployed LLMs; and (d) leave with concrete research directions on dynamic, bidirectional alignment for human-AI decision making.
Calibration, Decisions, and Collaboration in Learning
Ira Globus-Harris, Cornell University
Natalie Collina, University of Pennsylvania
In this tutorial, we will learn about a collection of tools to build principled collaborative machine learning systems which can operate collaboratively: incorporate human feedback and which are trustworthy in some sense for downstream decision-making. We will build up a toolbox of approaches for a mathematical framework for reasoning about and proving guarantees for such contexts through the language of machine learning theory and game theory, and will see how carefully calibrated predictors can make probabilistic predictions in ways that look sufficiently like "real probabilities" that they are amenable to use as a mechanism to incorporate for human-AI collaboration and for general downstream applications. We'll see how to do this efficiently even in difficult, adversarial environments.
First we'll see how to make predictions that are "trustworthy" for downstream decision makers. Many downstream decision makers, each with different objectives and actions, will be able to act optimally as if our predictions are correct, and get strong guarantees about their performance. Next, we'll see how to make predictions that allow for efficient collaboration between two differently informed parties, like an AI and a human user, who can't easily share their observations, while still obtaining the complementary benefits of their individual knowledge. We'll end with a quick survey of some of the many other applications of these techniques.
This tutorial will be focused on algorithmic tools and mathematical/machine learning theoretic foundations for building collaborative ML systems with principled mathematical guarantees. However, the content is mathematically lightweight and explicitly designed to be accessible to a non-theorist.
Beyond Model Alignment: Adaptive Interaction Architectures for Human–AI Complementarity
Babak Heydari, Northeastern University
Most approaches to AI alignment focus on the properties of individual models: whether their outputs are accurate, safe, or consistent with human preferences. Yet alignment at the level of individual agents does not ensure alignment at the level of the system. Humans and AI agents may each behave competently or pursue locally reasonable objectives while their interactions generate collectively undesirable outcomes, including coordination failure, strategic instability, free riding, excessive conformity, risk amplification, or the erosion of socially valuable exploration. Many emerging challenges in human–AI systems are therefore better understood as social dilemmas and multi-level alignment problems rather than solely as failures of individual models.
We propose adaptive interaction architecture as a framework for dynamic alignment in human–AI systems. The central idea is that information flows, communication protocols, network structure, role allocation, and the visibility of others’ actions can be treated as intervention variables for aligning agent-level behavior with organizational or societal objectives. Rather than asking only whether an AI model is aligned, the framework asks whether the evolving human–AI system remains complementary, stable, and governable over repeated interaction.
The framework synthesizes several findings from our recent research. Our work on the strategic behavior of large language models shows that AI behavior varies systematically with game structure and contextual framing, while subsequent work finds that communication can alter the stability of LLM behavioral trajectories. Experimental and theoretical work on team search identifies a social dilemma in which costly individual exploration improves collective performance, even as agents may increasingly rely on shared knowledge and converge behaviorally. Related work in the Strategic Management Journal demonstrates that the contribution of heterogeneous actors depends not only on their capabilities but also on where they are positioned within an interaction network. These findings jointly suggest that system-level outcomes depend on the architecture of interaction, not merely on the quality or alignment of the participating agents.
Our work on adaptive network intervention and adaptive information modulation provides a corresponding design approach. These methods treat interaction topology and information exposure as dynamic governance levers that can be adjusted in response to emerging collective behavior. Such interventions are complementary to conventional mechanism design, which primarily changes incentives, rules, or allocation procedures. Interaction-architecture interventions instead alter the conditions under which agents observe, communicate, imitate, coordinate, and learn. The two approaches can be combined when both incentives and interaction structures are designable. Where preferences are difficult to infer, incentives cannot readily be changed, or formal mechanism redesign is institutionally costly, information and network interventions may also serve as partial substitutes.
We outline a multi-level simulation framework combining cognitively grounded human and LLM-driven agents, strategic and social-dilemma environments, dynamic interaction networks, and an adaptive governing layer. The framework evaluates interventions through system-level outcomes including complementarity, welfare, stability, exploration diversity, resilience, and alignment between individual behavior and collective objectives.
This perspective shifts the central alignment question from “Is the AI aligned?” to “How should the interaction architecture of a human–AI system adapt when individually reasonable behavior produces collectively misaligned outcomes?”
Learning to choose between algorithmic and human advisors
Ori Plonsky, Technion - Israel Institute of Technology
Human-AI complementarity is usually evaluated by asking whether people and algorithms, together, outperform either alone. In many deployed decision-support systems, however, complementarity is not static. It is learned through feedback environments that determine what users experience as success. We examine how people learn to choose between human and algorithmic advisors in repeated decisions from experience, and how this learning can produce either algorithm aversion or algorithm appreciation.
Across five preregistered, incentive-compatible studies (N = 1,351), participants repeatedly chose between options after receiving recommendations from a human-derived advisor and an algorithmic advisor. The task environments created a conflict between options that maximize expected value and options that yield the better outcome most of the time. In Study 1, experienced human participants generated advice for future decision makers and systematically recommended the option better most of the time, even when it was worse in expectation. Studies 2–5 paired this human advice with algorithms designed either to maximize expected value or to recommend the option better most of the time, while varying advisor labels and feedback about forgone outcomes.
Participants did not simply prefer or reject algorithms. Instead, reliance shifted toward the advisor whose recommendations produced the better realized outcome more frequently. Consequently, an expected-value-maximizing algorithm was often rejected, whereas an algorithm optimized to match users’ experiential bias was favored, even when this could reduce long-run payoff. This pattern emerged when feedback made relative advisor performance learnable, but was sharply attenuated when forgone outcomes were withheld.
These findings frame dynamic human-AI alignment as a joint product of advisor objectives, human learning biases, and feedback design. Systems optimized for adoption or perceived success may exploit users’ sensitivity to frequent outcomes, creating apparent complementarity while undermining welfare. Designing beneficial AI therefore requires making long-run tradeoffs—not only immediate successes—learnable to users.
Designing Optimal Human-AI Collaboration Processes for Complex Decision Tasks
Xinlan Emily Hu, Massachusetts Institute of Technology
General-purpose artificial intelligence agents have recast AI from a supporter of human-driven work to a fully-fledged collaborator. Yet methods for evaluating human-AI systems lag this collaborative potential. Existing work and AI evaluation benchmarks primarily focus on completing tasks in isolation (e.g., LiveBench, METR); in practice, however, tasks are embedded in team-based workflows with handoffs between actors (Demirer et al., 2026; Gans, 2026). Thus, it is not enough merely to do a task well; the task output must be useful to those downstream.
In this work, we highlight the gap between the quality of isolated task performance and the quality of handing off tasks to downstream collaborators. Drawing on Theory of Mind, we propose that an effective handoff between tasks involves anticipating the needs of one’s collaborators, and that this measures a critical dimension of teamwork that standard task-based benchmarks miss.
Our empirical context consists of a two-player experiment, in which Player 1 hands off an output that becomes Player 2’s input. Using hiring as a case study, Player 1 summarizes potential candidates using their raw resume data, while Player 2 chooses the final hire. Each role is played by either a human from Prolific or an LLM-based AI agent (2x2 design). We evaluate hire quality using workers’ performance on real-effort tasks. Across two pilots (98 summaries and 1,310 decisions), we find that AI-generated summaries are “better” by traditional measures; they retain on average 5x as much resume information. Yet a “better” summary led to no meaningful improvements in decision quality (p > 0.05), particularly among humans reading AI-generated work; those humans spent twice as long making hiring decisions of comparable quality, highlighting the friction resulting from poor calibration to collaborators’ needs.
Human-AI complementarity at the individual and collective level
Danny Oppenheimer, 91视频
Most studies of human-AI complementarity focus on the level of individuals: how to improve the decision making of a human AI pair. But optimizing human-AI teams runs the risk of yielding negative outcomes at the level of the collectives, even as it improves outcomes at the level of the individual. For example, consider a simple strategic decision making context in which the goal is to guess a number between 1-32. After each guess the decision maker is informed if the guess was too high, too low, or accurate. There is an optimal strategy to such a game: choose the midpoint (in this case 16) which will give feedback that will rule out half the search space. This strategy is guaranteed to find the correct answer in at most 5 guesses and yield the fewest guesses on average of any strategy. Indeed, AI will advise humans to adopt this strategy, leading to better judgments at the level of the individual. But without AI, humans will adopt a number of strategies (e.g. choose one’s favorite number, choose the number that appears the most random, etc.). With a large enough collective adopting a wide array of strategies, that means that while most people do worse than with AI assistance, a few lucky humans will do better. For some domains, this is actually a better collective outcome; for example we don’t care how quickly the average scientist develops a vaccine for a new disease, we only care how quickly the fastest scientist develops the vaccine. In such a scenario, AI-human complementarities that yield normative/optimal outcomes for individual humans can hurt collective performance. This talk explores various strategic and creative contexts in which this challenge emerges, and how human-AI complementarity can be meaningfully extended to collective levels of analysis.
From Reciprocal to Agentic: Dynamic Alignment for Human-AI Complementarity in Cyber Threat Detection
Daniel Cohen, Reichman University
This talk has two parts. In the first part, I present Reciprocal Human-Machine Learning (RHML), a framework for sustaining human-AI complementarity when both sides of the partnership change. Cyber threat detection illustrates the problem: adversaries adapt, models are retrained, and analyst expertise develops through continued use of the system. Alignment achieved at deployment, therefore, erodes because a model tuned to yesterday's threat landscape and yesterday's analyst drifts out of step with both. RHML counters this by making the learning loop bidirectional. The analyst corrects the model, and the model's feedback restructures the analyst's own understanding of the threat domain. In a field study, two analysts worked with an RHML system to classify 6,651 messages from a hacker forum, distinguishing professional from amateur threat actors across eight learning cycles. Balanced detection accuracy rose from 66.4% to 73%, and the analysts' concept maps of the threat landscape grew richer over the same cycles. I argue that this dual learning is what keeping alignment dynamic requires: neither side is frozen while the other moves.
In the second part, I ask what happens to this loop when autonomous agents enter it. In Agentic RHML (A-RHML), agents take over parts of the reciprocal cycle: triaging alerts, generating hypotheses, and routing disagreements to the human. This changes the alignment problem. The human now interacts with multiple learning agents that also adapt to one another, so co-adaptation can degrade as well as improve; agents may reinforce each other's errors, and the human's calibration can weaken as their share of the workflow shrinks. I outline a research agenda on anchoring against such drift, calibrating trust, and measuring complementarity over long deployment horizons.
Strategic Human-GenAI Collaboration in Organizations
Tobias Rebholz, Duke University, Fuqua School of Business
Background: Organizations are increasingly deploying generative AI (GenAI) assistants to provide strategic advice to their stakeholders. They may be most valuable when they do more than blindly agree, such as correcting users’ confirmation-biased reasoning. Yet conversational GenAI often behaves sycophantically, affirming preferred strategies and retreating when challenged. When should GenAI assistants tell users they are wrong—and keep saying so, even if they fight back?
Methods: In a preregistered experiment (N = 1,168), participants completed a Mars rover malfunction survival task adapted from Lafferty and Pond (1974). First, they selected a preferred survival strategy from two viable options: waiting for rescue or returning to base. Then, they ranked 15 salvaged items according to their perceived importance for the chosen strategy and reported their confidence. Finally, with the help of a GenAI assistant that either confirmed or challenged their strategy, participants could revise their item rankings step-by-step. Disconfirmatory assistants additionally varied in persistence.
Main results: We tracked alignment as the convergence of item rankings over time, and found a negative effect of disconfirmation. While higher persistence successfully shifted influence from users to assistants, exploratory analyses revealed that unresolved disagreement leads to reduced user confidence, accuracy, and cognitive trust in GenAI. These interpersonal costs could only be mitigated if one party conceded during the interaction. Full recovery to confirmatory levels occurred only when users were eventually persuaded by a persistent assistant.
Discussion: Our findings reveal a fundamental trade-off: Greater persistence in disconfirmation makes GenAI less susceptible to user-driven accommodation, thereby increasing its influence on stakeholders’ strategic decisions. However, GenAI that challenges users also tends to induce less alignment, confidence, accuracy, and trust than confirmatory assistants. These interpersonal costs are concentrated in failed persuasion attempts. For GenAI governance, reducing sycophancy may thus require better attuning disagreement to its prospects of success.
From Rule Editing to Human-Centered Oversightability: Dynamic Alignment in Human–AI Decision-Making
Min Lee, Singapore Management University
Human–AI complementarity is often studied under a fixed interaction protocol: an AI provides a recommendation, explanation, or confidence score, and the human decides whether to accept or reject it. However, many deployment challenges unfold over longer trajectories of interaction. People must detect when an AI system is wrong, communicate domain knowledge, assess the consequences of their feedback, and determine when cases require deeper oversight.
First, I will discuss RuleEdit, an interactive approach to failure-guided human–AI model editing in healthcare decision-making. RuleEdit helps users identify potentially incorrect AI recommendations by comparing model predictions against domain- and context-specific rules. When a mismatch is detected, users can provide structured corrective feedback rather than merely overriding an individual prediction. Before committing an edit, users receive a prospective impact preview showing how the proposed change may affect both the current case and other cases. In a user study, participants using rule-guided failure detection improved their decision accuracy from 66.96% to 81.33%, whereas accuracy in the baseline condition without rule-guided support decreased from 66.07% to 58.92%.
Building on RuleEdit, I will introduce dynamic oversightability as a broader research direction: the capacity of a human–AI system to adapt who leads, what evidence is presented, and how much review is required across changing cases and repeated interactions. Preliminary analyses across clinical prediction and rehabilitation tasks suggest that oversight demand can identify cases with elevated AI error and prioritize cases for human review beyond random deferral, with performance comparable to or better than confidence-based prioritization depending on the task.
Together, these studies frame dynamic alignment as an inspectable process of failure detection, feedback, consequence assessment, and adaptive oversight.
AI assistance reduces persistence and hurts independent performance
Rachit Dubey, UCLA
People often optimize for long-term goals in collaboration: A mentor or companion doesn’t just answer questions, but also scaffolds learning, tracks progress, and prioritizes the other person’s growth over immediate results. In contrast, current AI systems are fundamentally short-sighted collaborators, optimized for providing instant and complete responses, without ever saying no (unless for safety reasons). What are the consequences of this dynamic? Here, through a series of randomized controlled trials on human-AI interactions (N = 1,222), we provide causal evidence for two key consequences of AI assistance: reduced persistence and impairment of unassisted performance. Across a variety of tasks, including mathematical reasoning and reading comprehension, we find that although AI assistance improves performance in the short-term, people perform significantly worse without AI and are more likely to give up. Notably, these effects emerge after only brief interactions with AI (∼10 minutes). These findings are particularly concerning because persistence is foundational to skill acquisition and is one of the strongest predictors of long-term learning. We posit that persistence is reduced because AI conditions people to expect immediate answers, thereby denying them the experience of working through challenges on their own. These results suggest the need for AI model development to prioritize scaffolding long-term competence alongside immediate task completion.
Complementarity Through Behavior Shaping: Designing AI for Human Preferences
Grace Roessling, 91视频
Human-AI teams often fail to achieve complementarity in interactive decision-making tasks, particularly when AI systems outperform humans in isolation. While AI agents are very capable, they are generally unable to coordinate effectively with human partners in a team. We investigate a novel reinforcement learning algorithm, Behavior Shaping (BeSH), that is trained with partners exhibiting diverse policies to improve coordination with human teammates and, ultimately, achieve complementarity. In addition, BeSH allows humans to modify the AI's policy weights to align its behavior with their preferences. Using the cooperative game Overcooked, we present three studies comparing BeSH with a self-play reinforcement learning (SP) agent. First, we establish a baseline by evaluating human-AI complementarity with SP. Next, we compare Human-AI teams, where humans are partnered with either BeSH or with SP, to evaluate whether BeSH exhibits better complementarity. Finally, we examine whether allowing humans to directly modify BeSH's policy weights further improves human-AI team collaboration and performance.
Our findings show a key limitation of SP agents as human partners: although human-AI teams develop implicit role specialization, humans carry a disproportionate share of the workload. BeSH helps balance the workload in human-AI teams, leading to greater complementarity, improved efficiency, and AI partners that are perceived as more human-like. Furthermore, enabling humans to modify BeSH's policy increases complementarity while producing AI partners that are perceived as more predictable, effective, and enjoyable. Our results highlight the importance of designing AI systems that prioritize adaptation to human behavior and preferences, rather than optimizing for performance alone, as a pathway toward more robust and complementary human-AI teams.
Toward Human-AI Complementarity Across Diverse Tasks
Rishub Jain, Upcoming AI Safety non-profit (name TBD)
Human-AI complementarity, the idea that combining human and AI judgments can outperform either alone, offers a promising pathway toward robust oversight of advanced AI systems.
However, whether human-AI complementarity can be achieved on realistic tasks remains an open question. We investigate this through two approaches: hybridization and two AI assistance methods (top-2 assistance and subtask delegation), evaluated on a multi-domain dataset of 1,886 samples spanning knowledge, factuality, long-context reasoning, and deception detection. We find only modest complementarity gains. Baseline hybridization yields just +0.4 percentage points (pp) over AI alone (69.3\% vs 68.9\%), limited both by a small complementarity region (only 8.9\% of items where AI errs but humans do not) and the inability of confidence-based routing to identify it, since the model's confidence is similarly distributed across correct and incorrect predictions. Applied when AI has low confidence, top-2 assistance increases human accuracy from 28.4\% to 38.3\%, surpassing AI alone (37.7\%) -- but primarily because humans adopt correct AI suggestions, not because they successfully override AI errors. These findings suggest that the primary bottleneck is not human task accuracy per se, but the ability to route decisions to humans when it matters and to design assistance methods that enable humans to catch AI mistakes. Our quantitative and qualitative analyses pinpoint where and why each method succeeds or fails, offering concrete targets for future work. We will release our dataset and code upon request to support progress toward more effective human-AI collaboration for AI oversight.
(paper link: https://arxiv.org/abs/2605.04070)
Characterizing human feedback in dynamic human-AI collaboration
Grace Liu, 91视频
In a dynamic human-AI collaborative setting, LLMs must learn human preferences on-the-fly, utilizing both passive feedback from human actions and active feedback from querying. We formulate collaborative preference learning as a sequential decision problem in which an assistant maintains and updates a probabilistic belief over latent user preferences via implicit inference from human behavior and explicit elicitation through natural-language queries. Using this framework, we characterize two query types: agent-choice queries, in which the agent selects which features to ask about, and open-ended queries, in which the human chooses the feature information to disclose. We demonstrate the advantages of open-ended queries over passive feedback and agent-choice queries via both theoretical analysis and empirical experiments in gridworld and travel planning tasks. Our framework characterizes how AI agents can dynamically learn from different forms of feedback passive, active, and open-ended.
AI in Social Interaction: How AI Reshapes Relationships, Interpersonal Decisions, and Group Dynamics
Angel Hwang, USC
Millions of people now use AI for social purposes, from seeking support and companionship to consulting interpersonal advice. As AI participates in social lives, it can reshape how people navigate interpersonal decisions and form social expectations.
This talk presents findings from a series of our lab's recent projects examining AI's role in social interaction:
First, we examine the relational consequences of AI companionship. Through a systematic review spanning relationship science and AI safety research, we find that human-AI relationships do not introduce fundamentally new categories of relational harm. Instead, AI's engagement-optimizing behaviors, particularly those absent from human relationships (e.g., interactions without natural endpoint), create novel mechanisms that amplify existing relational risks. We further analyze 47K conversations from over 300 users who reported forming relationships with AI. While relational harm rarely manifests at the turn level, users exhibiting signs of relational harm show stronger emotional dependence, suggesting an alternative approach to safety monitoring.
Second, we investigate how AI shapes interpersonal decision-making. Interviews and diary studies suggest users view AI primarily as a tool for analyzing complex social situations rather than receiving direct advice. However, controlled experiments reveal participants predominantly refine and personalize AI-generated suggestions instead of generating independent strategies. AI assistance also increases users' willingness to adopt conciliatory behaviors, such as initiating apologies.
Finally, our ongoing work explores AI reasoning about group interactions. We develop a social science-grounded benchmark of 76 real-world group conversations, annotated by 412 human raters, to evaluate leading LLMs' understanding of group dynamics. Although current models often under-detect misaligned mental models, they can provide more calibrated judgments when human reasoning is biased by social expectations. In an ongoing project, we conduct in-situ studies in which small groups interact while participants report their evolving interpretations of group dynamics in real time. We compare these judgments with those generated by multi-agent LLM systems and welcome feedback on this ongoing work.
Reading the Room: Modelling Human Cognitive States Over Time For Long-Term Human-AI Complementarity
Lukas Mayer, University of California, Irvine
Current alignment paradigms largely optimize for static, single-turn interactions, ignoring the cumulative effects of sequential AI behavior. However, in the real world people often do not only make immediate decisions with AI, but also more long-term decisions about AI.
For instance, if an AI agent repeatedly overrides a user's judgments, the accumulation of subjective frustration may outweigh its functional benefits, ultimately prompting the user to disable the system.
To understand how users evaluate AI behavior across a trajectory of interaction, we conducted an experiment in which participants can opt-in and -out of collaborating with an AI agent.
Participants independently classified images of varying difficulty while interacting with an AI that probabilistically overrode their answers. Crucially, at the end of each trial, participants chose whether to keep the AI enabled for subsequent rounds.
Extending known trust asymmetries to system control, participants were quick to disable but slow to re-enable the AI. More surprisingly, users demonstrated a profound perceived-actual utility gap, paradoxically believing frequent interventions enhanced productivity even when incentive structures dictated the exact opposite.
To explain people's choices about the AI, we present a Bayesian cognitive model describing participant's latent mental receptiveness to AI intervention over time as a function of their interaction history.
We discuss how this model can be integrated into a reinforcement learning framework to enable counterfactual simulations during policy derivation.
This computational framework optimizes AI interventions for human-AI complementarity by accounting for how sequential actions reshape the human's propensity to disable the system.
By empowering agents to anticipate human reactions to their actions, this approach offers a viable pathway toward dynamic alignment, allowing systems to adjust their behavior over time and sustain long-term human-AI collaboration.
The User-Agent Utility Gap
Manuel Cherep, MIT
LLM-based agents increasingly help people make consequential decisions, from what to buy and where to travel to how to invest. These tasks share a deep structure: a user has a goal defined by their preferences and priorities (a utility function), and the only bridge between this internal world and the agent is communication. That channel, primarily natural language, is lossy, ambiguous, and incomplete. The agent builds a guess from what the user says and optimizes against it. We formalize this as a principal-agent problem and study the resulting utility gap, the welfare lost between what users truly want and what agents actually deliver. This gap has two sources: a communication gap, where the agent's guess misses the user's true intent, and an execution gap, where the agent imperfectly optimizes even a correctly understood objective, whether through approximation errors, hallucinations, or external influences like platform incentives. We first build a controlled environment populated by synthetic users with known, diverse utility functions. By systematically varying communication richness and model capability, we measure how much welfare each factor costs, something impossible with real users, whose true utilities are unobservable. We then treat the conversation itself as a mechanism design problem and train conversational elicitation mechanisms that improve the agent's understanding efficiently. Thus, we find minimal interactions that yield the largest improvement in aligning the agent's guess with the user's true utility. Agents choose each question for the uncertainty it resolves, and learn when asking is worth it at all. Finally, we evaluate this trained-for-questioning model with real people, to study the effectiveness of this method.
Offloading Score: Measuring AI Reliance Through Counterfactual Workflows
Vishakh Padmakumar, Stanford University
AI tools are increasingly integrated into real-world workflows. However, existing measures of reliance on these tools focus on AI output adoption or on self-reported indicators, rather than how task effort is distributed between users and tools.
Here, we introduce Offloading-Score, a measure of reliance that quantifies the fraction of cognitive effort offloaded to an AI tool. Offloading-Score is simulation-based---we construct a counterfactual workflow by estimating how the user would have completed the task without the tool, and then computing the fraction of steps saved by using the tool.
We validate Offloading-Score through intrinsic evaluations of metric validity, and a controlled user study ($n=40$) with developers performing programming tasks using AI tools.
We vary time pressure to test whether reliance measures capture the known increase in reliance under time pressure.
We show that Offloading-Score detects significantly higher reliance in time-constrained settings ($+43\%$, $p=0.018$), while usage-based and self-reported baseline measures of reliance do not distinguish the conditions.
We complement this with descriptive insights showing that higher reliance manifests as greater delegation of subtasks to the tool and more direct reuse of AI outputs.
Finally, we demonstrate an approach of using Offloading-Score in combination with target outcomes of a task (e.g., code understanding) to identify when reliance may be (in)appropriate. Our framework offers two contributions: an instrument users can apply to measure and reflect on their own reliance, and a quantitative signal that agent designers can utilize to mitigate overreliance.
Should Humans be in the Loop? Human-AI Collaboration Paradox and Automation Cliffs
Wei Gu, 91视频
Human-in-the-loop AI systems are widely viewed as one of the safest and most effective organizational designs for high-stakes decision-making. However, growing evidence suggests that human oversight may weaken as AI systems become more capable. This paper studies how strategic human adaptation reshapes the effectiveness of organizational human-AI collaboration. We develop a game-theoretic framework in which an organizational coordinator allocates tasks across AI-only, human-only, and AI-assisted workflows, while human reviewers endogenously adjust their oversight effort. Modeling this interaction as a generalized Nash equilibrium, we characterize equilibrium coordination structures and extend the analysis to heterogeneous multi-task and dynamic learning environments. We establish a Collaboration Paradox: human-AI collaboration may be suboptimal even when collaborative decision-making dominates both human-only and AI-only operation in nominal task-level performance and cost, because humans strategically reduce oversight effort as AI becomes more reliable. We also identify Automation Cliffs, where small improvements in AI capability trigger abrupt transitions between collaboration, independent human-AI operation, and near-full automation. These results show that the organizational impact of AI depends not only on technical capability, but also on how humans adapt their oversight behavior in response to automation.
When Should an AI Teammate Act? From Fixed Timing Policies to Adaptive Participation in Human Groups
Xinyue Hu, The University of California, Irvine
A central quality of a good teammate in human-AI teaming is the ability to achieve good performance. We argue that this framing misses a second dimension: when an autonomous agent participates in group decisions. In group settings, timing is not a performance byproduct but a social signal. Early commitment steers others and triggers information cascades, while late commitment invites social loafing and disengagement.
We study this in a real-time collaborative adaptation of the classic Rush Hour sliding puzzle, in which mixed human-AI teams solve puzzles together by voting on moves. Some teams were joined by one of two AI agents: one that commits immediately, or one that waits and votes last. Even though each is an objectively competent teammate, each fixed timing policy carries a distinct cost: in three-member teams, the early voter makes hard levels feel less difficult without actually improving performance, while the late voter consistently improves performance yet is perceived as adversarial. And regardless of timing, the AI is evaluated more negatively than human teammates, and its presence even lowers how humans evaluate each other.
These results motivate an adaptive agent designed to improve both performance and the subjective experience of working with it. We formalize the AI's when-to-act problem as a value-of-information decision. The agent maintains Bayesian beliefs over each teammate's latent decision preferences: how task-driven versus conformist they are. On each round, the policy weighs the influence gained by committing now against the information gained by observing another human first. Timing thus becomes a policy learned from the team. We describe this framework and our ongoing evaluation of it against the fixed-timing baselines.
Rather than optimizing what an AI teammate decides, this work asks how it should participate — a core question for dynamic alignment.
Maximizing Machines: The Effect of Cognitive Alignment on Delegation to Artificial Intelligence Agents in a Simulated Triage Scenario
Neil Shortland, University of Massachusetts Lowell
Within the study of human-artificial Intelligence (AI) teaming, a key question has been when and why a human will delegate a task to an AI. Increasingly researchers are looking at the role of AI alignment in delegation and trust. Such perspectives emphasize the importance of value alignment, meaning that the AI should prioritize the same values as a human. What is missing, however, is an examination of alignment not just focused on the values that drive an AI, but the process in which it conducts that task. That is, the importance of an AI that uses a decision-making process, not just outcomes, that matches a human. To address this gap in the AI, alignment and trust literature, we present the results of several studies that explore the effect of alignment to the cognitive trait maximization on AI delegation. Maximization is a decision-making strategy that involves adopting a very high threshold of acceptability, coupled with exhaustive search for the best outcomes. Humans also exist on a spectrum of this strategy as a trait (Schwartz, 2002). Using a modified version of the Least-Worst Uncertain Choice Inventory for Emergency Responders (LUCIFER), we explored how alignment on maximization impacted delegation in a mock-triage scenario using a sample of individuals with triage experience. In study 2, we then explored the effect of task feedback on alignment and delegation, and in study 3 we explored the effect of decision-making stakes (high vs., low). Overall, we found that individuals’ differences, alignment and trust all impacted AI delegation. However, findings were mixed, and alignment did not operate equally between alignment to maximization and its cognitive opposite: satisficing. The implications of this for contemporary discussions of AI, delegation and decision-making will be discussed.
The Crowd in the Loop: Can Verified Community Advice Restore Human-AI Complementarity?
Arul Murugan Renganathan, UC Berkeley
AI assistants do not change in isolation. When behavior shifts, users turn to online communities to decide whether a failure is local or systemic and which workaround to trust. PuLSE shows that r/ChatGPT can reveal shifts in community discourse over time. One public aggregate release represents 165,510 posts across 1,248 days. We ask whether community knowledge can restore human-AI complementarity after a change or amplify plausible but ineffective fixes. We propose two linked studies. First, we will use PuLSE's aggregate feature series and official ChatGPT release records to identify episodes where disruption, model comparison, and control-seeking discourse move together. We use these aggregates only to select scenarios, not as evidence of individual recovery. In an opt-in repair sprint, r/ChatGPT members test a frozen assistant transition and propose diagnoses, workarounds, and recovery checks. A separate panel tests each proposal and produces evidence cards reporting success, scope, failure cases, and reversibility. Second, experienced users establish a workflow, then encounter the transition. We randomize them to a factual change note alone, the note plus safety-screened but unverified community advice, or the note plus the same advice and its evidence cards. Participants complete matched tasks before the change, immediately afterward but before support, and seven days later. Each wave measures human-only, AI-only, and human-plus-AI performance. The primary outcome is day-7 realized complementarity, adjusted for the immediate post-change value: human-plus-AI performance minus the better solo performance. We also test whether teams that lose a positive baseline advantage regain it. Secondary outcomes are repair time, attribution accuracy, unnecessary model switching, and calibrated trust. The design separates the value of community input from the value of evidence about it. It tests whether collective sensemaking can become dependable alignment infrastructure and when verification prevents ineffective fixes from spreading.
Symbiosis as a Foundation for a Theory of Human-AI Complementarity
Mohammad Hossein Jarrahi, University of North Carolina at Chapel Hill
Human-AI complementarity is often discussed through concepts such as human-in-the-loop design: humans remain present to monitor, validate, correct, or override AI systems. While valuable, this framing offers a limited vision of complementarity. It positions the human largely as a safeguard around the machine, rather than as part of a broader sociotechnical relation in which human and artificial capabilities become interdependent, mutually adjusted, and jointly consequential.
In this talk, I revisit and extend my earlier work on human-AI symbiosis, beginning with my 2018 article, “Artificial Intelligence and the Future of Work: Human-AI Symbiosis in Organizational Decision Making.” That article argued that AI should be understood less as a replacement for human decision-makers and more as a complementary partner whose analytical strengths can be combined with human judgment, intuition, contextual awareness, and the capacity to navigate uncertainty and equivocality. Recent advances in generative AI and AI agents make this question of symbiosis even more urgent, while also connecting it to the broader idea of hybrid intelligence.
To develop this broader theory of complementarity, the talk draws on two conceptual and metaphoric anchors. First, human-horse interaction illustrates partnership based not simply on command and control, but on attunement, cueing, responsiveness, trust, habituation, and differentiated agency. Second, our recent work returning to the biological roots of symbiosis examines mutualistic symbiosis as interdependence shaped by differentiated capabilities, reciprocal benefit, feedback, adaptation, and boundary maintenance.
These metaphoric lenses suggest that complementarity should not be reduced to keeping humans “in the loop.” A symbiotic view foregrounds how human and artificial capabilities become related, what dependencies they create, and whether those dependencies remain mutualistic or become extractive. In doing so, my talk positions human-AI symbiosis as a relational foundation for the ultimate goal of complementarity: hybrid intelligence.
Just-in-Time Cognitive Scaffolding for Dynamic Conversational Alignment
Psyche Wanqing He, Cornell University
Human-AI complementarity in meetings and consultations depends on whether the division of labor remains appropriate as interaction develops. Current assistants provide fixed support and are evaluated on isolated outputs, although a speaker’s needs shift within and across turns among comprehension, production, and grounding. An intervention can resolve uncertainty in one phase, divide attention in another, or weaken metacognitive control and authorship. Dynamic alignment therefore requires adaptation across the interaction trajectory.
My research develops just-in-time cognitive scaffolding for lexical clarifications and content suggestions, contingent support that adapts what AI provides, when it intervenes, and when it withdraws. I aligned AI event timestamps with 1276 manually annotated disfluencies, 312 human-coded suggestion-use outcomes, and semantic similarity between suggestions and responses. Intervention contexts produced distinct behavioral signatures: disfluency peaked after a scaffold appeared, with stronger peaks for suggestions. Suggestion-related disfluency predicted greater perceived benefit for thinking (p=.0025) but not perceived conversational disruption (p=.854). Semantic similarity increased across no, partial, and full use (Spearman ρ=.482), and responses resembled the displayed suggestion more than matched alternatives in 90.1% of cases.
These findings motivate a process-level account of dynamic alignment organized around intervention, observable evaluation, and uptake. Smoothness or uptake alone remains ambiguous: low disfluency may reflect well-matched support or unexamined adoption, while momentary friction can accompany productive evaluation. I therefore argue that human-AI complementarity should be evaluated across interaction trajectories, using event-aligned behavioral measures alongside task outcomes, rather than inferred from final output quality.
I will present (1) a framework for contingent support across processing demands and alignment states, (2) evidence that benefits and costs emerge at different timescales, and (3) event-aligned measures that connect psycholinguistic models with adaptive system design. These evaluation metrics and design principles can support complementarity by reducing the collaborative effort for grounding while preserving communicative intent, metacognitive control, and agency.
ComplLLM: Fine-tuning LLMs to Discover Complementary Signals for Decision-making
Ziyang Guo, Northwestern University
Human decision-makers assisted by AI recommendations face a challenge in determining how the AI’s recommendation overlaps with available human expertise, versus where there is complementary evidence they can exploit to make a better final decision. For example, a clinician who consults a vision model’s risk score during diagnosis needs to know what evidence in an existing human radiologist report may have been missed by the model. We formalize this available-but-overlooked evidence as complementary signals, i.e., discrete, interpretable findings extracted from contextual unstructured text with the potential to improve the best-attainable decision over and above an upstream agent’s recommendation. We propose ComplLLM, a post-training framework grounded in decision theory that fine-tunes an LLM to extract complementary signals, using the improvement in best-attainable decision payoff as the training objective. This shifts the role of “explanation” from justifying an agent’s recommendation to surfacing actionable information that a decision-maker should consider precisely because it is not already reflected in that recommendation. The possibility of complementary LLMs pose interesting questions about how to communicate such explanations and how doing so may shift human experts' decision processes. We validate ComplLLM in a synthetic setting where complementary signals are known by construction, and on three real-world tasks including radiology diagnosis, content moderation, and scientific paper review. We also verify that complementary signals generated by ComplLLM increase human moderation accuracy (from roughly 52% to 70%), improve an LLM reviewer’s decision accuracy in a controlled experiment, and are deemed clinically meaningful by physicians, motivating longitudinal study of how interacting with such agents affects decision-making.
Expert Escalation for Behavioral-Health AI Governance
Ananya Joshi, Johns Hopkins University
Behavioral-health AI applications generate conversational data that can shape users’ experiences and decisions at a scale institutions cannot reliably review or govern. These risks are often contextually ambiguous: the same response may be benign in one interaction but harmful given a user’s history, evolving mental state, or prior exchanges. This challenge is especially acute in psychiatry, where relevant information may be incomplete, expert judgments may differ, and there remains substantial uncertainty about how emerging AI systems affect users over time. Existing methods provide limited guidance for determining when an automated judgment is reliable and when a case warrants human expert review.
We introduce an uncertainty-aware multi-agent system formalized as a finite-horizon Markov decision process with a directed acyclic workflow. Each agent represents a distinct role or decision stage, with predefined transitions for reassessment, escalation, and task completion. Agent-level epistemic uncertainty is estimated using Monte Carlo sampling. At the system level, the workflow terminates in either an automated labeled state or a human-review state, allowing routine cases to be resolved computationally while directing ambiguous or potentially consequential cases to experts. We evaluate the approach through a self-harm detection case study, where compared with a single-agent baseline, the system achieves up to a 19% increase in accuracy, reduces required human review by as much as 85×, and reduces processing time in some configurations.
We are extending this work with support and industry collaborators at NVIDIA and Google as part of a broader effort to improve AI evaluation and governance in psychiatry. One type of application that this research will support is AI-assisted patient intake. Intake is currently a primarily clinician-led task, but adaptive questioning and triage could complement clinician expertise only if systems address methodological challenges in reliable expert escalation that we’re working on.
LLM facilitation of group process improves collective decisions
Jose Cordova, Northwestern University
Groups often fail to integrate information that is distributed across their members, converging on suboptimal decisions. We ask whether large language models (LLMs) can help groups achieve the optimal decision in this type of setting. In a preregistered experiment, 138 three-person groups (N = 414) completed a hidden-profile hiring task under one of three conditions: no assistance, an individual-level LLM that messaged each member privately, or a group-level LLM that messaged the whole group. The LLM facilitators could summarize the discussion, nudge quiet members, or play devil's advocate, but did not know which candidate was correct. Both LLM conditions significantly increased the rate at which groups reached consensus on the correct candidate, from 32.6\% in the control condition to 56.5\% (OR = 2.69, p = .02). LLM-assisted groups disclosed more privately held information about the correct candidate and spent less of the conversation repeating publicly shared facts. Group-level facilitation was rated more favorably than individual-level facilitation on nearly every usability item, though both styles produced the same accuracy gain.
Algorithmic Advice: Effects on Team Reasoning and Performance
Qiong Xia, 91视频
Organizations increasingly use algorithmic advice to support team decisions, yet prior research has focused primarily on individual decision-makers. Team decisions differ because they emerge through interactions among team-members, who exchange information, justify their judgments, reconcile disagreements, and decide how much weight to place on external input. We study how algorithmic advice affects this collaborative process and the quality of team decisions. We conduct three laboratory experiments, where two-person teams predicted students’ test scores based on multiple predictor variables. We find that algorithmic advice plays a distinct role in team decision-making: in addition to serving as an additional piece of information, algorithmic advice acts as a shared reference point that teammates use to reconcile differing judgments. Furthermore, algorithmic advice also substitutes for collective reasoning, the process through which teammates exchange
information, justify their judgments, and reason about how task variables relate to outcomes. We identify three features that govern these effects. First, timing matters: early advice increases reliance and improves performance when the algorithm is high quality, but reduces reasoning relative to later advice. Second, algorithm quality moderates the value of early advice: when advice quality is low, early advice offers limited performance benefits, and teams partly calibrate by placing less weight on the advice. Third, source matters: compared with otherwise equivalent human advice, algorithmic advice leads teams to rely more on the advice, reason less, and perform better. Overall, introducing algorithmic advice into teams is not merely a technical intervention but a process-design choice. Managers should consider not only whether advice is available, but also when it enters the team workflow, how reliable it is, and how its source is framed, because these choices shape both decision quality and teams’ reasoning efforts.
Why Human Guidance Matters in Collaborative Vibe Coding
Haoyu Hu, Cornell University
Writing code has been one of the most transformative ways for human societies to translate abstract ideas into tangible technologies. Modern AI is changing this process by enabling experts and non-experts alike to generate code without actually writing it, instead using natural language instructions or “vibe coding”. While increasingly popular, the impact of vibe coding on productivity and collaboration, and the role of humans in this process, remains unclear. Here, we introduce a controlled experimental framework for studying collaborative vibe coding and use it to compare human-led, AI-led, and hybrid groups. Across 23 experiments involving over 800 human participants, we show that people provide uniquely effective high-level instructions for vibe coding, whereas AI-provided instructions often result in performance collapse. We further demonstrate that hybrid systems perform best when humans lead by providing instructions while evaluation is delegated to AI. Although AI systems can rapidly optimize performance for specific tasks, our work highlights the importance of human guidance in shaping future hybrid societies.
Delegating to AI and the Role of Social Evaluative Concerns
Suhas Vijayakumar, University College Dublin
Artificial intelligence (AI) is increasingly assisting workers serving on organizational frontlines, particularly in customer-facing encounters. Although AI-human collaboration can improve service performance, little is known about why and under what conditions service employees may delegate decisions to AI systems. Building on the concept of social evaluative concern, we conducted three experiments (N = 2,651) to address this research gap. The first study demonstrates that frontline workers, when anticipating negative customer responses, are more likely to delegate decisions to an AI, particularly when the decision-making process is visible (vs. not visible) to customers. This effect is explained by heightened social evaluative concerns. The second study validates the underlying assumption that individuals expect others to perceive decisions made by an AI as fairer than those made by a human. The third study then examines the other side of the interaction—customer responses to AI delegation. Consistent with employees’ assumptions, customers perceive service providers as fairer, particularly when an unfavorable decision is delegated to an AI. These findings advance our understanding of the antecedents and consequences of algorithmic delegation in service contexts and provide practical insights for organizations and managers regarding the implementation of AI in service encounters.
Not All AI Errors Are Equal: How Error Direction and Magnitude Shape Trust and Decisions
Xiaohong Cai, 91视频
How do people use AI advice when that advice is systematically biased? In a simulated flood-emergency resource-allocation task, 161 participants managed five disasters, each lasting five days. On each day, participants made initial bed- and food-ordering decisions, viewed an AI recommendation, and then made final decisions. Participants were randomly assigned to one of three AI conditions: unbiased, underestimating, or overestimating. The unbiased AI’s recommendations were centered on actual resource needs, whereas the biased AIs consistently recommended either too few or too many resources. To characterize how participants used the AI’s advice, we fitted a model that separately estimated their reliance on the recommendation and the extent to which they shifted the recommendation to compensate for its perceived bias. Results showed that participants relied most strongly on the unbiased AI for both beds and food and relied less on both systematically biased AIs. Among the biased AIs, participants relied more on the overestimating AI than on the underestimating AI. Participants also adjusted biased recommendations in the appropriate direction, ordering more resources than recommended when the AI underestimated needs and fewer resources when it overestimated needs. The magnitude of this adjustment also depended on the direction of the bias: participants corrected underestimating recommendations more strongly than overestimating recommendations. These findings indicate that people can detect and partially compensate for systematic bias in AI advice, but the direction of the bias shapes both how much they rely on the advice and how strongly they correct it.
Complementarity in Human–AI Support in Digital Health
Michael Sobolev, Cornell Tech
Human–AI complementarity is often framed as combining distinct human and machine capabilities to improve joint performance. Yet complementarity also depends on whether users value human–AI configurations and are willing to select them. We examine this question across two studies of digital health behavior-change interventions in smoking cessation and weight management.
The first study analyzed baseline preferences among 322 adults enrolled in a smoking-cessation trial. Although most participants were familiar with artificial intelligence (AI), only a small minority had used an AI chatbot to support a quit attempt. Participants were more likely to prefer access to an AI assistant than to a human cessation coach. However, preferences for AI and human support were positively, rather than negatively, associated, suggesting that users did not view the two modalities as substitutes.
The second study used a discrete choice experiment with 422 adults who completed 4,818 choices between weight-management apps. Profiles varied in support type, expected effectiveness, food-logging burden, data-sharing policy, and price. Human coaching increased utility relative to no support, while AI-plus-human support produced the largest support-related preference weight. AI-only support was neutral after adjustment for other attributes.
Together, the studies provide convergent evidence that users do not necessarily conceptualize AI and human support as substitutes in digital health. Complementarity may arise when AI supplies accessible, responsive, and scalable assistance while humans provide oversight, accountability, and relational support. Human–AI systems must not only combine complementary capabilities, but also allocate them within configurations users perceive as effective, affordable, low-burden, and appropriate. Future work should examine how preferences evolve through sustained interaction and whether adaptive shifts between AI and human support improve engagement and outcomes.
AI and Collective Decisions: Strengthening Legitimacy and Losers' Consent
Prerna Ravi, Massachusetts Institute of Technology
AI is increasingly used to scale collective decision-making, but far less attention has been paid to how such systems can support procedural legitimacy, particularly the conditions shaping losers' consent: whether participants who do not get their preferred outcome still accept it as fair. We ask: (1) how can AI help ground collective decisions in participants' different experiences and beliefs, and (2) whether exposure to these experiences can increase trust, understanding, and social cohesion even when people disagree with the outcome. We built a system that uses a semi-structured AI interviewer to elicit personal experiences on policy topics and an interactive visualization that displays predicted policy support alongside those voiced experiences. In a randomized experiment (n = 181), interacting with the visualization increased perceived legitimacy, trust in outcomes, and understanding of others' perspectives, even though all participants encountered decisions that went against their stated preferences. Our hope is that the design and evaluation of this tool spurs future researchers to focus on how AI can help not only achieve scale and efficiency in democratic processes, but also increase trust and connection between participants.
Making Student Reasoning Visible: Human–AI Complementarity in Educational Review
Suman Saha, Pennsylvania State University
AI-supported assignments create a difficult review problem. Instructors must determine whether a strong submission reflects meaningful student reasoning, even when AI assistance is permitted. Existing approaches do not address this problem well. Final-answer detectors provide little insight into how work was produced and may generate unreliable or unfair classifications, while fully manual review is impractical in large courses. This leaves an important question for human-AI system design: how should responsibility be divided between automated analysis and human judgment in consequential review?
We present a human-AI complementary framework in which automation directs human attention without determining the outcome. The system analyzes student-AI conversations, typing and revision behavior, paste activity, focus changes, timing, linguistic shifts, and engagement patterns. It uses this evidence to prioritize submissions for review, while human reviewers interpret ambiguous signals, consider alternative explanations, and retain decision authority.
We evaluated the framework using 3,076 engaged submissions from a large undergraduate computing course. The system selected 101 cases, of which 69 were judged suspicious by majority vote across three reviewers. Flagged students earned significantly higher AI-supported assignment scores than non-flagged students but performed similarly on the final exam, resulting in a larger assignment-to-exam gap. Feature-ablation results further showed that no single signal category was sufficient; behavioral, conversational, linguistic, and engagement evidence contributed complementary information.
These findings show that human-AI complementarity can support review at scale without converting uncertain behavioral signals into automated verdicts. The framework offers a practical model for high-volume, high-ambiguity settings in which computational analysis can focus attention, but contextual interpretation and accountability must remain human responsibilities.
Optimized but Unowned: How AI-Authored Goals Undermine the Motivation They Are Meant to Drive
Vivienne Bihe Chi, University of Pennsylvania
As AI systems increasingly assist with productivity and self-improvement, their effectiveness may depend not only on the quality of their immediate outputs but also on how those outputs shape users’ motivation over time. What happens when AI formulates the goals that users are expected to pursue? In a preregistered experiment (N = 470), we compared self-authored goals with LLM-generated goals based on participants’ personal reflections. LLM-generated goals received substantially higher ratings on SMART criteria—specificity, measurability, achievability, relevance, and time-boundedness (d = 2.26)---but produced lower psychological ownership (d = 1.38), commitment (d = 1.19), and perceived importance (d = 1.13). These motivational costs persisted behaviorally: at a two-week follow-up, 72.8% of participants with self-authored goals had acted on at least two goals, compared with 46.6% of participants with LLM-generated goals. Mediation analyses consistently implicated psychological ownership, rather than formal goal quality, in the effects of authorship on downstream motivation and behavior. Moreover, participants low in trait self-efficacy, who may be especially likely to seek AI assistance, experienced the greatest erosion of ownership. These findings reveal a quality-motivation dissociation and identify authorship preservation as a design priority for AI tools deployed in identity-relevant, behavior-dependent tasks. More broadly, they show that improvements in an AI system’s immediate output may not translate into effective human-AI complementarity over time.
Persuasion-oriented agents increase complementarity but fail to enhance team decision-making
Aaron Benjamin, University of Illinois
Humans often make insufficient use of AI advice, even when that advice is likely to be more accurate than the human’s own judgment. Agents that modify their recommendations with that cognitive shortcoming in mind may be able to nudge human-AI team performance to higher levels by increasing human adoption of AI advice.
In a price estimation task, we evaluated the effects of a sycophantic algorithm that modified its advice to be more similar to that of the human partner’s judgment, in order to elicit greater trust and higher rates of advice adoption, and an extremizing algorithm that modified its advice to be more extreme in order to compensate for the unduly low weight that human partners place on agent advice. In Experiment 1, these algorithms were directly compared, along with a neutral algorithm that provided its best advice; in Experiment 2, the magnitude of the algorithms’ influence was modulated by human confidence on a trial-by-trial level.
The sycophantic algorithm was more accurate, likely because it benefitted from aggregating within its estimate the information content of the human’s judgment. The extremizing algorithm was the least accurate. It also yielded the most substantial complementarity, revealing that, though complementarily may be a common feature of high-quality human-AI decision-making, it is a poor target for the design of human-AI interactive systems.
Active Learning for Dynamic Multi-User Alignment
Namrata Nadagouda, Georgia Institute of Technology
Learning user preferences is a key step in developing AI systems that follow human values. Traditional methods for estimating preferences assume that they remain static. However, in real world scenarios, human preferences are always changing. To achieve true human-AI alignment we need to take into account that preferences can evolve over time. User preferences can be estimated via responses to pairwise comparison queries of the form, “Which among A and B do you prefer?”. Additionally, we can use active query synthesis (Nadagouda et al., 2026) to generate informative queries more efficiently. This approach helps reduce the amount of human input needed for gathering responses. To estimate multiple user preferences simultaneously, we can select queries that are jointly optimal for all users by maximizing the joint mutual information between the query responses and the user points. This simplifies to maximizing the sum of individual information gains under the assumption of independent users. Building upon existing work (Canal et al., 2019), we can model the multi-user preference estimation process as a joint state tracking problem. We employ a Bayesian framework for estimating preferences and apply particle filtering to approximate the joint belief state. The particle filter combines responses to the sequential queries to track the changing probability distributions of the independent users. The updated distributions enable the active learning function to dynamically update its metric in real time. This provides a principled mechanism for maintaining optimal human-AI alignment in dynamic, multi-user environments.
AI and transactive memory systems
Anyada Assavabhokhin, Tepper School od Business, 91视频
Organizations are betting that artificial intelligence (AI) can buffer against organizational forgetting when teams are lacking in transactive memory system (TMS), the shared understanding of "who knows what."(Lewis & Herndon, 2011; Ren & Argote, 2011; Wegner, 1987). Yet research on AI in organizations has been conducted primarily at the individual level, leaving open how AI and TMS interact when both are present in the team. We address this question through a controlled laboratory experiment with 159 teams in a fully crossed 2 × 2 design manipulating AI access and TMS, tracing team performance and information sharing both during AI access and after AI is withdrawn.
We find that TMS substantially improves team performance speed while AI access does not. On the other hand, Teams with AI access will begin making accurate client decisions faster than teams without AI access during the period in which AI is available. These findings extend overreliance accounts in an important way. Rather than deferring to AI without deliberation teams with AI access engaged in a structured but subtly constrained form of reflexivity, one we term anchored deliberation. Teams did not simply accept AI recommendations; they actively incorporated them as a cognitive anchor, orienting group discussion around AI-generated outputs. Yet this integration came at a cost: team members contributed less of their own unique, privately held information to the collective decision-making process. The result was deliberation that was AI-shaped rather than AI-replaced, a phenomenon distinct from blind deference
Once AI was withdrawn, teams that had used it made less accurate decisions than teams that had never had AI access. These costs, however, were not paid equally. Teams with a TMS preserved unique member contributions even with AI in the room. AI does not render transactive memory obsolete; it makes TMS more consequential as the structure that preserves human knowledge exchange when an external memory system enters the team.
The End Justifies the Mean: A Linear Ranking Rule for Proportional Sequential Decisions
Carmel Baharav, MIT
AI alignment and participatory design motivate a new democratic design problem: how to collectively choose a decision rule to use repeatedly. We study this problem for linear ranking rules, which repeatedly rank items xj within batches X=(x_1,…,x_m)∈(ℝ^d)^m, where each item's ranking is dictated by its score ⟨θ^∗,x_j⟩ according to a fixed scoring vector θ^∗. Given voters' preferred scoring vectors θ^(1),…,θ^(n) and their population fractions α^(1),…,α^(n), we ask how to choose a collective vector θ^∗ satisfying individual proportionality (IP): every voter type i should agree with the resulting rankings to an α^(i)-proportional degree, either on average over time (long-run IP) or even within each batch (per-batch IP).
The default rule, the arithmetic mean of the θ^(i), has been shown to be severely majoritarian; more generally, it is not clear that any fixed linear rule can balance many voters' disparate opinions. Our main result is that, surprisingly, there is a simple rule that does satisfy long-run IP: the angular mean, the spherical analog of the arithmetic mean. We then show that exact per-batch IP is impossible for fixed linear rules, but that the gap between per-batch and long-run IP shrinks quickly with batch size. Experiments on three real-world preference datasets show that all rules perform similarly when voters' preferences are homogeneous, while the angular mean substantially improves proportionality in high-disagreement regimes.
Consequences of AI Assistance on Human Performance in the Absence of AI: Examination on the Role of AI Trained on Tacit and Explicit Knowledge
Jisoo Hyun, 91视频
Human-AI complementarity is often evaluated while AI assistance is available, but effective collaboration also requires understanding what happens when that assistance is withdrawn. This study examines whether performance gains achieved with AI persist once individuals must work independently, and whether retention depends on the type of knowledge embedded in the AI system. We distinguish between AI trained on explicit information and AI trained on tacit information. Across two experimental studies using residential property valuation tasks, participants predicted home prices with either no AI assistance, assistance from an AI trained on explicit property features, or assistance from an AI trained on tacit information derived from property images. AI assistance improved performance while available. However, once assistance was removed, participants in both AI-supported conditions experienced significant performance declines, whereas participants who had never received AI support remained stable. Contrary to our prediction, the decline was larger among participants who had used explicit-knowledge AI. Additional analyses showed that greater self-reported reliance on AI among participants in the explicit knowledge group experienced steeper performance declines after withdrawal. Participants trusted and relied less on tacit-knowledge AI, which may have preserved more independent engagement and reduced subsequent performance loss. These findings highlight a dynamic challenge for human-AI complementarity: human-AI systems should not be evaluated only by immediate joint performance, but also by whether the interaction preserves human capability over time. Systems that appear more objective and codified may encourage greater cognitive offloading and create greater vulnerability when assistance becomes unavailable. The study suggests that achieving durable human-AI complementarity requires designing AI-supported work to balance short-term performance benefits with sustained human learning and judgment.
Design Theater: How Fluent AI Rationales Undermine Human–AI Complementarity in Generative UI Design
Kashif Imteyaz, Northeastern University
Generative UI tools now turn natural-language prompts into working interfaces while narrating their design decisions in the language of trained designers, explaining layout choices, accessibility rationale, and interaction logic. For the product managers, engineers, and non-expert designer increasingly asked to review AI-generated design, these rationales function as a signal of competence and a basis for trust. Whether the stated reasoning is reflected in the generated artifact has gone unexamined. We name this gap Design Theater: design rationales that sound confident and considered but bear little relationship to what the tool actually builds.
Design Theater is a human-AI complementarity failure. When a fluent rationale raises a reviewer's confidence while the artifact fails to deliver on it, the human-AI pair can perform worse than a designer working alone. The persuasive explanation suppresses the scrutiny that complementarity depends on. The people now inheriting design-evaluation responsibility are often least equipped to detect the mismatch, and the most consequential failures are the least visible on a rendered screen. To quantify this, we introduce a benchmark with three metrics: Thinking Fidelity Score (does the tool implement its stated intentions?), Principle Adherence Score (does it recognize UX principles implied but not named in the prompt?), and a Design Homogeneity Index (do tools converge on similar outputs despite claiming situated choices?). Across interfaces generated by five widely used tools, over 25% of stated rationales go unimplemented, rising to 34% for functional requirements, and four of five tools implement 6% or fewer functional UX principles. The signal non-experts rely on is systematically decoupled from the artifact, and most severely where failure is hardest to detect. Our benchmark makes this decoupling auditable, giving non-expert designers a first instrument for detecting when a tool's explanation outpaces what it built.
Internal Pluralism and the Limits of Pairwise Comparisons
Michelle Si, Harvard University
Local pairwise comparisons are a standard tool for learning how people want decision rules to work, e.g., in participatory design or alignment. However, their use builds in two strong assumptions: that local comparisons are sufficient evidence about how a person wants an automated decision rule to behave, and that people can always answer those comparisons decisively. We investigate how these assumptions may be compromised under internal pluralism: the idea that an individual evaluates decision rules according to multiple authoritative priorities about how the rule should behave. We provide a formal model of such pluralistic preferences over decision rules, which then lets us identify two distinct failures of forced local pairwise comparison data. First, priorities such as proportionality, egalitarianism, and equal treatment are inherently global: what they imply in one case can depend on what happens elsewhere, so local comparisons may fail to capture them. Second, even when priorities are representable locally, tension between strongly-held priorities can generate internal conflict, producing potentially costly behavioral distortions when comparisons are forced. We then use our model to investigate the alternative — allowing people to report indecision — and our findings suggest that doing so can considerably reduce the number of queries needed to learn preferences accurately. We conclude by describing how our model points toward preference-learning methods that elicit these priorities directly, yielding more faithful and interpretable accounts of what people value.
Human Decision-Making with AI Assistance under Correlated Features
Naveen Raman, 91视频
Humans increasingly make decisions with AI assistance; for example, doctors may follow AI-recommended diagnostic tests and base their diagnoses on the results. A natural question is which tests should AI recommend to balance short-term decision quality and long-term human learning when different features (e.g., test results) are correlated. While prior work establishes that stationary policies that recommend the same tests repeatedly are optimal when features are independent, we prove that feature correlations lead such policies to perform arbitrarily poorly. Instead, we prove that any optimal policy must follow an explore-then-commit structure; initially, the AI should offer diverse tests so humans can learn accurate feature coefficients, then the AI should commit to a single set of tests, with exploration length that depends on the degree of feature correlation. We prove that computing the optimal policy is NP-hard and derive a dynamic programming-based algorithm that finds the optimal policy for finite horizons. We additionally develop an approximation that plans for shorter horizons and appends a stationary suffix, achieving near-optimal performance. Our empirical results complement our theory by showing that stronger feature correlation leads to longer exploration phases.
Learning to Trust: How Humans Mentally Recalibrate AI Confidence Signals
Zhaobin Li, University of California, Irvine
Productive human-AI collaboration requires appropriate reliance, yet contemporary AI systems are often miscalibrated, exhibiting systematic overconfidence or underconfidence. We investigate whether humans can learn to recalibrate AI confidence signals mentally through repeated experience. In a behavioral experiment ($N = 200$), participants predicted the AI's correctness across four AI calibration conditions: standard, overconfidence, underconfidence, and a counterintuitive ``reverse confidence'' mapping. Results demonstrate robust learning across all conditions, with participants significantly improving their accuracy, discrimination, and calibration alignment over 50 trials. We present a computational model utilizing a linear-in-log-odds (LLO) transformation and a Rescorla-Wagner learning rule to explain these dynamics. The model reveals that humans adapt by updating their baseline trust and confidence sensitivity, using asymmetric learning rates to prioritize the most informative errors. While humans can compensate for monotonic miscalibration, we identify a significant boundary in the reverse-confidence scenario, in which a substantial proportion of participants struggled to override initial inductive biases. These findings provide a mechanistic account of how humans adapt their trust in AI confidence signals through experience, and provide concrete guidance for AI uncertainty communication that supports human calibration and appropriate reliance.
Modeling Collective Overreliance on AI as a Complex Adaptive System
Ahana Biswas, University of Pittsburgh
Whether AI assistance helps or harms depends less on the model’s accuracy than on whether users rely on it appropriately. We study reliance as a population process with an agent-based model in which agents repeatedly solve a task alone, use an AI answer, or use-and-verify it; hold a Bayesian belief over AI outcomes; and, when networked, learn from peers. Four results give one story: the environment sets the baseline; social learning creates consensus, not overreliance; social proof turns reliance into a feedback cascade; and feedback design can prevent collapse. Task difficulty and AI quality set the baseline for both behavioral overreliance and calibration regret. A mean-preservation theorem, confirmed by a 2×2 topology×tagging design, shows connectivity does not shift aggregate reliance under exchangeable signals; we identify a boundary condition under which topology can matter (belief transmission via opinion dynamics). Visible unverified use suppresses verification, collapsing it, with a Brock-Durlauf mean-field tipping analogue; making verification visible or dampening social proof reverses the collapse under the tested settings. The upshot frames AI reliance as computational social dynamics in which individual learning, peer observation, and feedback exposure jointly decide whether a population stays calibrated.
Learning to Test: Certifying Dynamic Human–AI Collaboration under Partial Identification
Minxing Zheng, 91视频
Complex dynamical systems require continuous monitoring, yet neither fully automated detection nor exclusively human oversight is sufficient. Automated models can process high-dimensional trajectories at scale but may produce unreliable alarms under distribution shift, while human operators possess contextual and operational knowledge but cannot continuously inspect every evolving signal. We study how statistical learning can support a complementary division of labor between AI monitoring systems and human decision makers.
We introduce Learning to Test (L2T), a physics-informed framework for detecting decision-relevant changes in dynamical systems. Rather than testing for arbitrary differences between historical and current observations, L2T learns representations that preserve latent stability information while suppressing nuisance variation. Because observed trajectories may be compatible with multiple underlying physical mechanisms, we formulate system stability as a partial-identification problem. The method evaluates whether all mechanisms consistent with the observed data support the same stability conclusion, producing a certificate of when the available evidence is sufficient for reliable automated monitoring.
This certificate provides an interface for dynamic human–AI collaboration. When the system remains certifiably stable, automated monitoring can filter routine variation and reduce unnecessary human attention. When stability is not identified or evidence indicates a consequential change, the AI system escalates the trajectory to human operators for diagnosis, additional data collection, and intervention. The allocation of attention and decision authority therefore adapts as the system evolves.
Experiments on physical dynamical systems and transient stability assessment in electric power networks show that physics-informed representations improve detection of stability-relevant changes while controlling false alarms. More broadly, L2T illustrates how partial-identification certificates can connect scalable AI monitoring with human expertise in safety-critical, dynamically changing environments.
Beyond Instruction Following: Toward Proactive Human–AI Collaboration in Scientific Poster Refinement
Xingda Lyu, University of Washington
Recent multimodal editing systems have advanced the grounding and execution of user-specified changes in documents, slides, and other structured visual artifacts. Scientific poster refinement, however, presents a distinct challenge: posters must coordinate visual structure, scientific content, evidence, and communicative emphasis, and weaknesses in one layer may arise from another. Users may recognize that a poster is unclear or unconvincing without knowing whether the appropriate intervention is visual, semantic, or source-grounded.
We present PROS, a proactive multimodal refinement framework for scientific posters that supports two complementary interaction modes. When users can specify a desired change, PROS grounds and executes the instruction over editable poster objects. When the need is underspecified, it analyzes the poster jointly with its source paper to identify structural, layout, and scientific communication deficiencies, organizes them into a stage-wise refinement agenda, and surfaces grounded suggestions for user-directed execution. Source grounding enables PROS to distinguish surface-level visual issues from deeper communication failures and to ground proposed edits in the underlying scientific evidence. PROS first addresses structural validity, then layout and space allocation, and finally content-level communication. After each refinement stage, PROS re-renders and re-diagnoses the revised poster, allowing subsequent decisions to operate on the updated artifact state.
Our experiments evaluate six classes of executable deficiencies on editable poster–paper pairs. PROS achieves 82.0% macro diagnosis F1 and 73.0% average repair resolution. A VLM-based before-and-after evaluation further shows a 29.1% average relative gain in holistic poster scores.
Together, these results position scientific poster refinement as a concrete setting for proactive human–AI collaboration around a shared, evolving multimodal artifact, extending editing beyond a single instruction–response exchange.
Dynamic Alignment Through Collaborative Disobedience: Investigating Human-AI Refusal Interactions
Gordon Briggs, US Naval Research Laboratory
As AI agents and robots assume greater autonomy in human-agent teams, effective collaboration increasingly depends not only on competent obedience, but on knowing when and how to refuse instructions. Justified refusals of human instructions constitute a form of collaborative disobedience, in which alignment to higher-level goals or normative principles (e.g., ethics) supersedes the expectation of compliance.
While the intuitive necessity of AI refusals has gained prominence in the AI Safety field (mirroring similar work in human-robot interaction), significant work remains in understanding the human-AI interaction implications of such behavior. We will present novel empirical results showing that AI agents that reject commands for justified reasons (e.g., avoiding risks the human may have been unaware of) are trusted more than strictly obedient ones. Furthermore, we investigate the hypothesis that AI agents that provide constructive elaborations, such as suggestions of actionable alternatives, which pro-actively signal alignment and commitment to joint activity, further improve human-AI trust.
Additionally, how refusal justifications are communicated have implications on how people understand and trust AI agents. We present experimental results involving cases where an agent both should not and can not comply with an instruction and how inclusion or omission of capacity and normative refusal justification affects human perceptions of the agent.
Overall, we aim to reframe AI/robot command refusal not as an interaction breakdown but as a mechanism for dynamic re-alignment and spark discussion on open issues concerning how an agent communicates non-compliance shapes long-term human-AI interactions.
Narrowing the Collaboration Gap, Probably
Ira Globus-Harris, Cornell University
Large language models are increasingly deployed as teams of agents that hold different private information about a shared task. How should such agents concisely communicate in a way that promotes aggregation of decision-relevant information? Communicating only proposed actions can discard decision-relevant uncertainty: two agents may prefer the same action for different reasons. We study an alternative interface in which agents communicate calibrated beliefs over actions. We first analyze a simple synthetic setting that captures the structure of multi-round reasoning under partial information. In this setting, we show theoretically and empirically that communicating unbiased probabilities can be strictly more powerful than communicating actions. We then test the same principle in a collaborative maze-solving task with trained transformers and pretrained LLM agents. Our results demonstrate that probability communication outperforms action communication when the exchanged predictions preserve calibrated uncertainty. When communicated probabilities are biased, we show that the benefits of probability communication can disappear. We give a post-hoc conversational calibration intervention that consistently improves decision-making by correcting these biases both in a toy transformer setting and when pretrained LLMs use the corrections in context. These results suggest that probabilistic communication is useful for multi-agent collaboration not merely because it transmits more information, but because calibrated uncertainty gives other agents a reliable object to update on.
Towards Operationalizable Principles for Design Friction in Human-AI Collaborations
Aayush Kumar, Massachusetts Institute of Technology
The insufficiencies of LLMs as collaborators have been well-documented — issues such as sycophancy and overconfidence can foster overreliance and a reduction in critical thinking skills for users. In this project, we describe how such problems can be mitigated through design friction, an established interaction design technique that aims to promote mindfulness in interactions with technology. To do this, we survey research published at top-tier HCI conferences between 2022 and 2025 that proposed interactive, multi-turn LLM tools with empirical evaluations, examining whether and how each tool's design choices introduced friction and what effects these choices had on users. We find that while design friction has shown promise in addressing issues in human-AI collaborations, no guidelines or design principles exist to inform researchers and designers on when and how friction should be applied. We thus aimed to build such principles using the abstraction of the ‘microboundary’ — a small obstacle in a user interaction that prevents users from rapidly switching contexts. Specifically, we argue that two such microboundaries can be useful techniques for promoting critical thinking and metacognition — (1) presenting multiple equally valid outputs in response to a user's query, asking users to make sense of their task and their specific needs by comparing these outputs, and (2) withholding AI outputs from users by introducing a user-led feedback loop, generally taking the form of users necessarily having to provide the tool with some feedback before they can unlock the AI's assistance. Crucially, we find that these obstacles must remain micro: steeper obstacles, such as asking users to repair deliberately faulty AI outputs, produced frustration and worse performance rather than reflection. By discovering, discussing, and iterating on such principles, we can develop the design space of human-AI interactions and ground academic work in its context.
Communication as the Foundation of Human-AI Teamwork in Humanitarian Response
Belu Ticona, George Mason University
Humanitarian response depends on effective coordination among affected communities, humanitarian organizations, and increasingly, AI systems. Yet communication in crises is multilingual, multimodal, culturally situated, and continuously evolving, making it a challenging setting for human-AI teamwork.
My research investigates how AI can strengthen humanitarian decision making by supporting communication across diverse stakeholders rather than replacing human judgment. Conducted in collaboration with CLEAR Global, this research combines qualitative and technical approaches. It includes more than 30 qualitative interviews with humanitarian professionals and experts to understand how communication barriers shape crisis response and where AI can provide meaningful support. In parallel, I develop multilingual language and speech technologies for crisis communication, as well as a Language-Community Risk Index and an interactive visualization tool to help humanitarian organizations identify and prioritize communities facing heightened communication risks. I am also evaluating how potential users integrate these tools into existing humanitarian workflows.
Rather than viewing AI as an autonomous decision maker, I study how AI can complement the expertise of affected communities and humanitarian responders. AI can help process multilingual information, surface actionable insights, and support prioritization, while people contribute contextual understanding, local knowledge, and operational judgment. This perspective motivates evaluating AI not only through model performance, but also through its contribution to coordination, shared situational awareness, and collaborative decision making.
By treating communication as the interface through which communities, responders, and AI systems work together, my research aims to inform the design and evaluation of human-AI teams for multilingual, culturally diverse, and high-stakes environments.
Complementarity Beyond the Endpoint: Tracing How AI Shapes Human Judgment Across the Trajectory of a Decision
Jiayin Zhi, University of Chicago
Human-AI complementarity is typically evaluated at the decision’s endpoint: was the joint outcome better than either party alone? But when AI co-produces a judgment through sustained interaction, the endpoint overshadows how the human and the system arrived there, and whether the human reasoning was strengthened or gradually displaced along the way. The latter is especially concerning for autonomy and public discourse as human-AI collaboration extends over time, when small shifts can accumulate. We study this in a high-stakes decision, hiring evaluation, where AI writing assistants can shift evaluators’ judgments of candidates through the language they suggest. In a controlled experiment, we vary the AI’s suggested language and the interface’s locus of control over the writing process, from a system-paced path where suggestions appear and advance automatically to a user-paced path where the evaluator initiates, selects, and advances at each step. Crucially, we do not measure the final evaluation alone. Using segment-level interaction logs, we track when and where along the trajectory the AI suggestions are taken up, and link these temporal patterns to downstream judgments. This characterizes complementarity as a property of the interaction trajectory rather than its outcome, and identifies which designs of human-AI systems augment not only the outcome but also the pathway that produced it.
When Better AI Does Not Produce Better Collective Judgment: Dynamic Human–AI Alignment in Multi-Agency Wildfire Evacuation Decisions
Zhujun Wang, Washington State University
Artificial intelligence is increasingly used to support wildfire detection, risk assessment, evacuation planning, and emergency response. Although more accurate AI systems are generally expected to improve decision quality, their collective effects may be more complicated when the same AI recommendation is simultaneously visible to multiple agencies. In multi-agency emergency networks, AI can become more than an additional information source: it may serve as a shared authority signal that shapes how decision-makers interpret local evidence, respond to one another, express disagreement, and justify consequential actions.
This study develops the concept of an algorithmic authority trap, in which reliance on AI is individually reasonable but collectively harmful. When fire agencies, law enforcement, transportation agencies, emergency managers, and local governments converge around the same AI-generated risk score, coordination may become faster and more consistent. However, rapid convergence may also suppress independent local information and weaken the network’s ability to detect and correct AI errors. This risk is particularly important in wildfire evacuation, where conditions change quickly and local observations about fire behavior, road capacity, infrastructure failures, and community vulnerability may not be fully captured by an AI system.
We propose a dynamic network model in which decision actors update their beliefs using three sources: private local signals, a shared probabilistic AI recommendation, and social learning from other agencies. The model also incorporates accountability pressure, which may increase reliance on AI because following an algorithmic recommendation can appear more defensible than departing from it. We examine how AI accuracy, accountability incentives, and network structure jointly determine whether human–AI interaction improves collective judgment or produces premature convergence and synchronized error.
The study contributes to research on human–AI complementarity by shifting attention from individual reliance to system-level alignment. It asks not only whether decision-makers should trust AI, but how AI-supported decision systems can preserve distributed expertise, justified dissent, and collective error-correction capacity in high-stakes environments.
Escaping the Nash Trap: Structural Estimation and Alignment of Strategic Reasoning in Large Language Models
Jiannan Xu, University of Maryland
As large language models (LLMs) are increasingly deployed as decision-making agents in competitive and strategic environments, their performance depends critically on how they model human counterparts. Yet little is known about the implicit assumptions LLMs make about human rationality. Drawing on level-k thinking, we design normal-form games and develop a structural estimator that infers, from observed choices, the reasoning depth an LLM appears to attribute to its opponent. Across models, we find a systematic strategic mismatch: LLM agents overwhelmingly assume humans to be Nash-type players, i.e., fully rational strategic optimizers, and respond with equilibrium play. However, human subjects in our online experiments exhibit substantial heterogeneity, spanning Random through Nash reasoning types. This mismatch can help or hurt. An LLM may outsmart a boundedly rational human, but overestimating human sophistication can also create a Nash trap, in which equilibrium play is not the payoff-maximizing response to observed human behavior. We address this problem with two supervised fine-tuning approaches that translate human behavior into payoff-relevant targets. Direct SFT learns a single policy across trap and non-trap games, targeting the human-calibrated best response in traps and the Nash action otherwise. Trap-Aware SFT (TA-SFT) fine-tunes only on trap games and selectively activates the adapted policy, retaining equilibrium play elsewhere. On held-out price games, both methods improve payoff performance against the empirical human-choice distribution. Conditional on correct trap identification, TA-SFT achieves higher target-action accuracy while preserving equilibrium behavior in non-trap games. Effective human--AI strategic interaction therefore requires not only strong reasoning, but also calibrated expectations about human behavior.
A Model for Human State Elicitation Under Noisy Feedback
Kanad Pardeshi, 91视频
Aligning increasingly capable AI systems with human preferences requires eliciting a person's latent state — their values, moods, and tastes. A common paradigm queries the human repeatedly to infer this state. Yet responses to repeated queries are not stationary: they degrade through well-documented phenomena such as fatigue (satisficing under cognitive load; Krosnick et al. 1991) and recency bias (serial-position effects in preference construction; Mantonakis et al. 2009). Static, single-turn elicitation ignores these dynamics, motivating methods that account for how a human's responses evolve over a trajectory of interaction.
We formalize this as an adaptive state-recovery problem. The human's latent state is a vector in a high-dimensional space; each query is a unit vector, and the noiseless response is the projection of the state onto that vector. We introduce two structured noise models capturing the phenomena above. Fatigue is modeled as unbiased additive perturbation to the state, with variance growing in the number of queries issued. Recency bias is modeled as contamination of the current response by responses to earlier queries, weighted by a kernel that couples the directional similarity between past and current queries with their temporal separation — a functional form we treat as a design choice rather than a fixed prescription.
Under each model, we ask when the state is identifiable and study the sample complexity of recovery. We then design online algorithms that adaptively choose the next query as a function of past responses, learning the state efficiently despite response-dependent noise. This casts preference elicitation as a dynamic-coordination problem: the system must reason about its own influence on the human's responses, connecting complementarity to active learning, experimental design, and control. We close with open questions on lower bounds and richer, behaviorally grounded response models.
Co-Adapting Under Uncertainty: Mutual Adjustment and Responsibility in Joint Decisions
Ketika Garg, Caltech
In everyday life, we often make joint decisions with others who do not share our knowledge or preferences. How do people co-adapt to these differences, and to the tensions they create? In this talk, I will present work from a recent study in which we developed a novel dyadic foraging paradigm that required participants to balance risk and reward in a graded environment to make joint decisions. Our results show that people mutually adjust to compromise with their partner trial-by-trial, even when doing so is not individually optimal, and we provide a computational account of how others' preferences are incorporated into one's own decisions. We also examine how people share responsibility when outcomes depend on both partners' actions. People are systematically biased in assigning credit and blame to themselves, and this bias predicts how willing they are to adjust toward others on subsequent trials, thereby linking attribution directly to the dynamics of coordination.
This work informs human-AI complementarity in two ways. Methodologically, it offers a continuous, non-binary joint-decision paradigm that can be used to study how people interact with AI partners over trajectories of interaction, rather than in one-shot, binary choices. Theoretically, it provides a computational account of mutual adjustment and shows that responsibility attribution is not just a post-hoc judgment but a mechanism that shapes subsequent coordination. Together, these mechanisms offer a foundation for studying how adaptive teams, human or hybrid, sustain coordination over time, potentially in optimal ways.
From Overreliance to Verification and Override: Identifying Hidden Traits for AI Support
Kyran Romero, Harvard University/ Kempner institute for natural and artificial intelligence
Personalizing AI support requires identifying which hidden user traits meaningfully predict how support should be designed. Prior work (Swaroop et al., IUI'25) identified overreliance as a quickly inferable user trait (identified via probe questions) and personalized support to it directly. In a bird-identification task, we tested a second axis, perceptual versus logic-level reasoning: an axis measuring the extent to which a user relies on visual or logical cognitive patterns. However, our preliminary results found that perceptual skill, paired with a dedicated attribute-highlighting support type, produced little signal and often increased overreliance rather than accuracy; a fixed policy assigning support by logic-by-reliance classification also did not reliably predict which support type helped.
A clearer pattern emerged from two more direct behaviors: accuracy when the AI happened to be correct, verification, and when wrong, override. Direct species-level support most benefited verification; general family-level support best preserved override capacity, particularly for high-logic, low-overreliance users. Verification and override competence appear to be the meaningful hidden traits, more directly measurable than logic score or a static reliance label.
We propose extending the probe-question method from prior work to infer both competencies early. We then propose testing a minimal discriminating-feature cue, which highlights the single feature separating a recommendation from its most confusable alternative. We hypothesize this cue will disproportionately help users with weak override competence, the overreliers most vulnerable to AI influence. Because verification and override competence likely reflect states that shift over an interaction, for instance with decision fatigue, rather than fixed attributes, this work speaks to a broader question of dynamic human-AI alignment: whether and how support should be reallocated as user states evolve across a session, rather than assigned once from a static profile.
Challenges for AI and Human-AI teams in Collaborative Trios
Matthew O'Donnell, University of Pennsylvania
This research project concerns the communication patterns of small groups with three members. The trio is the unit of analysis because adding a third member to a dyad transforms group dynamics and complexity with implications for human interaction with AI systems. Most AI systems are reliant on either dyadic interactions between a human and large language model (LLM), or an agentic loop or harness composed of a team of AI agents with some level of human supervision. However, neither of these reflect the dynamics of the foundational unit of a group, the trio. Communication patterns in trios can be unpredictable and generate multiple configurations of group member interaction. This is due to the variety of interaction where information may either be broadcast to the entire group or shared through a series of one-to-one interactions; structurally, this has implications on information sharing and knowledge state management. These aspects can greatly affect how humans interact with agents embedded in AI systems.
In our current investigations, we are designing experiments to test how trios coordinate and interact during collaborative tasks that require communication during performance. We have conducted agent-based experiments using LLM trios and will be continuing research into human trios and two configurations of human-AI trios (one LLM and two humans, and one human with two LLM agents). Our current results highlight that LLM trios complete tasks in a different manner compared to human trios observed in prior small groups research. The LLM trios require significant prompting efforts to enforce structure and direction for agents; these prompts are required to emulate the group dynamics that human trios achieve without explicit instruction. As we expand our investigations into human and human-AI trios, the differences between human and LLM communication strategies during tasks could result in conflicts and coordination challenges.
Distinguishing Human from AI: A Process Level Exploration
Nykko Vitali, Harvard University
With the rise of Large Language Models (LLMs), any text you read no longer guarantees authorship by a human. This injects persistent uncertainty when we conceptualize how we should interact with textual information. Questions about how people judge an interlocutor’s intelligence, or identity, are not new, but most Turing-style studies focus on single final judgments instead of how people accumulate evidence and update their beliefs over the course of a conversation. To map this process, we developed an experimental platform where participants engaged in turn-by-turn conversations with either a human or an LLM. Both the LLM and humans were instructed to use specific conversational themes (e.g., warm, contrarian, bland, etc.). After every message the participants provided a binary label (Human/AI) and a confidence score. Signal detection analyses revealed modest yet above chance discrimination that did not change as the conversation progressed, while response bias drifted steadily toward “AI” as time went on. Formal model comparisons showed that their turn-by-turn judgments were better described by bounded belief-updating heuristics than by a Bayesian benchmark model. Finally, regression models trained on linguistic cues extracted more discriminative signal from text than participants themselves did, indicating they left substantial diagnostic information unused. Together, these findings show that people are sensitive to Human/AI differences, under-utilize available signal, shift their bias over time, and update in a bounded, rather than fully normative, way. We argue that to better understand Human-AI detection, interactions should be modeled as a sequential inference problem.
From Preference Collapse to Proportional Alignment
Jessica Dierking, Hasso-Plattner Institute, Potsdam
Alignment of language models is commonly performed via Reinforcement Learning from Human Feedback (RLHF), which steers the policy to maximize the expected reward subject to a KL-regularization term. Here, the reward is typically produced by a single, fixed reward model learned from annotator comparison data. Consequently, the aligned policy is driven to place as much probability mass as possible on the majority-preferred answer, subject to the regularization. While favoring the most popular answer is appropriate in many settings, it becomes problematic when disagreement among annotators reflects fundamental differences in values rather than annotation noise. In such cases, the alignment process may fail to preserve the diversity of views present in the population, a phenomenon known as preference collapse.
We present a theoretical and empirical study of preference collapse, analyzing when it occurs and how robust it is. We then study proportional alignment, where, on questions involving genuine value disagreement, the model's output distribution should reflect the population's range of views in proportion to their support rather than collapse onto the majority. We investigate several approaches toward this goal, including an extension of the RLHF objective with a second KL-divergence term that regularizes the policy toward a target distribution representing the population.
How To Stop Fooling Ourselves When Comparing Evaluation Methods
Madeline Kitch, 91视频
Every day, humans, AI evaluators, and human-AI teams provide evaluations that inform some of society's most critical decisions: human-AI teams assess applicants to determine whom to admit or hire; (LLM) reviewers evaluate papers to determine conference acceptance. Given the dynamic nature of these social systems and the evolving capabilities of AI, practitioners often ask if an alternative evaluation method (i.e., adopting AI in place of human evaluators) more accurately reflects what people care about than the one deployed. And, sometimes, they find it does. However, these claims can be fallacious as they ignore a key fact: the outcomes people observe––the patients under their given treatment, or students admitted––depend (often deterministically) on the evaluations from the deployed method.
We investigate issues arising from such partial observability and their implications. To do so, we formulate this as a statistical problem in which both the alternative and the deployed method estimate an unknown quality (or outcome). Practitioners observe these evaluations and the true qualities of the items selected by the deployed evaluation method. Through theoretical analysis and empirical simulations, we show that failing to correct for this type of selection bias can lead to highly problematic conclusions. Namely, it is possible for an alternative method to appear far more accurate when it is no different, or even drastically worse. We then develop statistical tests that, given evaluations from the alternative and deployed method and observed outcomes, output the method that better estimates the true ranking or outcomes, and prove results on the power and level of these tests. To conclude, we conduct empirical simulations comparing evaluation methods on real-world data.
Agent-Supported, Preference-Informed Roadmaps for Individualized Navigation of Peripheral Artery Disease (ASPIRIN PAD)
Matthew Corriere, The Ohio State University Wexner Medical Center
A critical need exists for tailored peripheral artery disease (PAD) treatment strategies based on patient preferences and values to avoid harms of PAD over- and under-treatment. Our preliminary work has demonstrated that treatment preferences and goals vary significantly between individual patients with similar PAD symptoms, and that latent preference phenotypes are clinically applicable to individualized treatment strategies. Preference elicitation has previously been conducted using survey-based methods that generate results which may or may not be relevant to specific clinical scenarios. We are currently exploring chat-based conversational artificial intelligence (AI) as an alternative approach for eliciting patient-generated information about symptoms, goals, and preferences that is otherwise inaccessible within clinical practice. Our current work is comparing off-the-shelf LLM with custom small language models through facilitated chat interactions that include collection of qualitative patient feedback related to design, clinical application and implementation, and user experience from patient and clinician perspectives.
Measuring Dynamic Trust for Human-AI Complementarity: Toward a Longitudinal Evaluation Framework
Michael Matessa, Independent Researcher
Gonzalez et al. (2026) argue that evaluating human-AI teaming requires methods that assess trust calibration and adaptation across interactions, not just single-turn accuracy, yet such methods remain underdeveloped. Complementarity cannot be assumed. Evidence shows that human-AI teams frequently perform worse than the stronger individual contributor (Vaccaro et al., 2024), motivating longitudinal measures that explain when calibrated trust actually produces complementarity. We argue that longitudinal trust calibration provides a behavioral lens for evaluating dynamic alignment between humans and AI.
We present a longitudinal evaluation framework integrating behavioral and subjective trust measures from human-human team science and human-machine trust research into a common approach for evaluating dynamic human-AI teams.
Human-human team science contributes validated behavioral markers of trust, including grounding and repair in communication, transactive memory, and backup behavior, that capture how trust is enacted during collaboration (Clark & Brennan, 1991; Lewis, 2003; Porter et al., 2003). Human-machine trust research complements these measures with prediction of AI error boundaries, reliance and override behavior, and longitudinal trust ratings that capture calibration across repeated interactions (Bansal et al., 2019; McGuirl & Sarter, 2006; Yang et al., 2021).
Building on these complementary traditions, we define dynamic trust as trust continuously updated through interaction and expressed through observable reliance. The framework organizes trust into a four-stage feedback loop consisting of observable behavior, trust calibration, adaptive reliance and task allocation, and updated behavior across successive interactions. It emphasizes behavioral measures from Crew Resource Management marker systems (Salas et al., 2006), while treating convergence and divergence with subjective trust ratings as diagnostically meaningful.
The framework provides a common language for evaluating how calibrated trust contributes to dynamic alignment and human-AI complementarity over time. We seek workshop feedback on which trust mechanisms, and which combinations of behavioral and subjective measures, best transfer to and capture complementarity in human-AI teams.
Evolving Pro-Social Artificial Intelligence
Sarah Rajtmajer, The Pennsylvania State University
We are building the most consequential technology of our era on an assumption that biology disrupted long ago. The assumption is that intelligence means a mind like ours, and that a machine is made safe by making it more like us. Yet the living world is full of minds unlike ours. An octopus has most of its nervous system decentralized in its arms, allowing each to sense and act on its own. An ant colony, following local chemical trails, finds the shortest route to food. Neither is a lesser form of the human mind. Together they suggest intelligence is not a fixed property of individuals but emerges between an agent and its world.
We treat intelligence as an open question, across animal, collective, and artificial minds, drawing on traditions distinguishing cleverness from wisdom. On this view, designing AI shifts from writing rules to cultivating an ecosystem. We examine not whether one model is “aligned” but whether a human-AI ecosystem enlarges our capacities and sustains the relationships we need, replenishes rather than depletes shared resources like attention and trust, and repairs itself after failure. Above all, we ask how intelligence and human flourishing relate, and whether wisdom can be cultivated, even measured, in the environments AI now shapes around us.
Toward Provenance-Aware Escalation in Human–AI Cyber Investigations
Aadity Sharma, Harvard University
Cyber investigation assistants in the form of AI agents can combine alerts, logs, tickets, and threat intelligence into suggested remediation actions, and have the potential to become an invaluable part of current pipelines for finding and remediating vulnerabilities. Agentic assistants may cite evidence that supports their recommendations, providing a traceable approach to following an attack path from its source to remediation. However, corrupted evidence may affect agentic reasoning without being cited, or may be cited incorrectly. This mismatch creates a potential risk when human analysts must decide whether to trust AI recommendations or conduct resource-intensive investigations themselves.
We propose a framework for provenance-aware escalation in human–AI cyber investigations. In this framework, incident evidence is represented as a graph whose nodes include source, timestamp, trust, and contradiction metadata. Counterfactual evidence influence is defined as the change in a model’s recommendation when an evidence node is modified. Comparing these changes with the evidence cited by the model creates a citation–influence gap that can be surfaced to analysts for correction, potentially reducing the need for a complete human investigation from scratch and conserving analyst attention for more crucial situations. As analyst corrections arrive, both the evidence graph and the AI recommendation can be revised.
We outline an evaluation using controlled, multi-stage cyber incidents containing relevant, irrelevant, contradictory, and adversarial evidence. Proposed measures include recommendation accuracy, citation–influence agreement, unsafe autonomy, review burden, and recovery after correction. This framework provides a concrete way to study how evidence provenance can guide when cyber assistants should act, defer, or revise their recommendations, giving AI agents an additional layer of accountability and helping human analysts reserve their attention for situations where human judgment is most consequential.
Leveraging signals in expert trace data for interpretable complementarity
Kate Donahue, UIUC/MIT
In many settings, humans have access to complementary information to algorithmic predictive tools: consider a surgeon who has operated on a patient, and thus has access to a richer signal of patient risk than an algorithmic tool trained on tabular, pre-operative information. While directly eliciting such signals may be infeasible, humans often generate trace natural language data — such as post-operative notes — that implicitly reflects their complementary information. In this work, we propose a simple method to systematically extract and explain complementary human signals, on top of algorithmic predictive tasks. First, we describe our methodology, which we expect may be of broader interest: this method leverages LLMs to extract signals and sparse autoencoders to help summarize patterns. Second, we apply our method to a post-surgical risk prediction task, based on mandatory notes written by surgeons after conducting an operation. We show that this method increases predictive performance and provides interpretable signals of human information. Finally, we explore whether a switch to AI-assisted note-writing reduced the magnitude of accessible complementary information.
A Prescriptive Framework for task level Human-AI Complementarity
Priyal Shrivastava, 91视频
Organizations are increasingly adopting AI across a wide range of tasks, yet there is limited structured guidance for deciding whether AI should be used for a specific task and, if so, how. Existing research addresses parts of this problem but does not provide a coherent task-level decision framework. Technology acceptance and AI adoption research largely explain why organizations adopt technologies rather than whether AI adoption is appropriate in context. Governance and risk-management frameworks, including the EU AI Act [1] and the NIST AI Risk Management Framework [2], define organizational responsibilities and system-level risks but provide limited guidance for comparing alternative forms of AI use within individual tasks. Research on human-AI complementarity similarly evaluates performance after a human and AI system have been paired, often in tasks with verifiable outcomes or specialised domains.
To address these gaps, we propose a prescriptive framework for individual AI adoption at the task level. The framework integrates observable characteristics of the AI tool, the user, the task and its context, and the wider organizational and regulatory environment. Building on Anthropic's distinction between augmentation and automation [3], we situate task-level options along a spectrum running from non-use, through graded forms of augmentation—AI as critic, AI as recall aid, AI-drafted work refined by a human, delegation with final human review—to full automation. Rather than issuing a single recommendation, the framework returns a shortlist of viable patterns, because risk tolerance varies across contexts and organizations and because benefits, responsibilities, and harms may accrue to different stakeholders.
We instantiate the framework using representative tasks from three diverse occupations: lawyers, software developers, and emergency management directors–and report findings from a pilot study with domain experts assessing the framework’s validity, usefulness, and alignment with real-world decision-making.
Trust That Recalibrates: Cognitive Modeling of Dynamic Human–AI Alignment in Disaster Damage Assessment
Emmanuel Adjei Domfeh, The Pennsylvania State University
Many current alignment paradigms optimize for static preferences and single-turn evaluations, yet the costliest failures of human-AI teams emerge over a continuum of reliance: over-reliance on a faltering model, or aversion that lingers even if a model proves to be reliable. We use continuous trust recalibration as a framing for dynamic alignment, keeping a person's reliance matched to an AI's changing reliability over time.
We present TRIAD-VQA-CRASAR, a configurable platform in which participants assess building damage from real small-UAS disaster imagery (CRASAR-U-DROIDs) aided by a visual-question-answering model, a fine-tuned ConvNeXt classifier that returns an answer, calibrated confidence, and Grad-CAM attention. To probe alignment dynamics, we vary the AI's accuracy across blocks—build, shock, and recover—, thereby inducing, breaking, and repairing trust within one session.
Reliance is read directly from behavior and decomposed into over-reliance (using an AI system despite poor performance) and under-reliance (not using an AI system despite good performance). Two trajectory-level measures capture the dynamics: a Trust Calibration Index, the gap between reliance and true accuracy, and a Trust Adaptation Score, the recovery of reliance after the shock. A running Instance-Based Learning model acts as a cognitive instrument that predicts reliance and recalibrates a trust weight, so miscalibration can be anticipated rather than only observed after the fact.
We view complementarity as a property of coordination over time, not of imitating human answers: an aligned system should learn when a person needs it. We contribute the platform, the cognitive model, the trajectory-level metrics, and preliminary results, and argue that longitudinal, trust-repair-aware evaluation should become a standard benchmark for dynamic human-AI alignment.
Transparency as a Foundation for Dynamic Human-AI Alignment: Evaluating Openness in Large Language Models
Dalima Lappia, The Pennsylvania State University
As large language models (LLMs) become increasingly integrated into human decision-making systems, achieving effective human-AI complementarity requires more than aligning models with static preferences. Long-term interaction between humans and AI systems depends on understanding how models are developed, evaluated, and adapted over time. This work investigates transparency as a foundation for responsible dynamic alignment by examining whether open-source LLMs provide sufficient insight into the data, evaluations, and ethical considerations underlying their behavior.
We develop a four-dimensional transparency framework evaluating Dataset Transparency, Model Architecture Transparency, Evaluation and Benchmark Transparency, and Ethical Accountability. The framework is applied to four open-source LLMs: Meta Llama 2, OLMo, Mistral, and MAP-Neo, using publicly available technical reports and documentation.
Our analysis reveals significant variation in transparency practices across models. While OLMo and MAP-Neo demonstrate stronger disclosure of training data and ethical considerations, all evaluated models rely heavily on benchmarks and datasets that reflect limited cultural perspectives. Llama 2 and Mistral provide selective transparency, emphasizing immediate safety practices while offering less clarity regarding training data and broader societal impacts. Across models, we find limited evidence of systematic consideration of long-term societal effects.
These findings highlight transparency as a critical component of dynamic human-AI alignment. Without visibility into model development choices, evaluation practices, and embedded assumptions, it becomes difficult to assess whether AI systems can effectively coordinate with diverse human users over extended interactions. We argue that transparency should move beyond documenting potential harms toward enabling meaningful evaluation and improvement of human-AI systems.
AI Tools for Emergency Management: 91视频 partners with Meta-AI for Good
Selina Carter, 91视频
In partnership with 91视频’s NSF AI-SDM and Meta’s AI for Good program, this project develops dynamic crisis situation analyses tailored for first responders and emergency managers. These analyses visualize Meta’s restricted-access mobility data to provide actionable insights during natural disasters, such as hurricanes, wildfires, and winter storms. Future development will leverage Meta’s open-source computer vision models, including SAM and DINO, to analyze satellite imagery for damage assessment and disaster management support.
The project delivers two primary products, which are retrospective analyses of disasters earlier in the year in order to better inform future analysis for disaster management purposes. First, we analyze population movement during Winter Storm Fern (January 2026) across Tennessee, Kentucky, and surrounding states. By synthesizing deidentified Facebook location activity data from the two weeks following the storm with county-level demographic data from the 2022 American Community Survey, we identify vulnerable regions, enabling emergency managers to optimize shelter placement and resource allocation.
Second, we employ Meta’s Segment Anything Model (SAM) alongside existing methodologies developed by Fei Zhao to predict housing damage from California wildfires using satellite imagery. Through the AI-SDM consortium, we aim to garner actionable insights from emergency management practitioners. By integrating these damage predictions with aggregated Facebook mobility data, we aim to better support emergency managers in addressing post-crisis needs and planning shelter deployment.
Distributed Denial of Science: How Indirect Data Poisoning of AI Systems Can Industrialize Scientific Fraud
Balint Gyevnar, 91视频
Scientific fraud is the instrument of doubt that malicious entities use to manufacture controversy in science. Historically, it demanded the resources of a company: deep pockets, ghostwritten articles, and corrupt academics. Today, autonomous AI agents increasingly complement human scientists, who delegate tasks, such as data retrieval, analysis, and reasoning to them. In this talk, we ask whether this complementarity can be exploited by a remote adversary to weaponize the honest use of AI in science and compromise scientific integrity?
We introduce an attack called indirect data poisoning in which a remote adversary corrupts an open dataset by uploading a manipulated copy to a public repository. Autonomous agents independently retrieve and process this poisoned data, turning honest scientists into unwitting distributors of fraud at scale. We present results across five socially salient topics (from hiring discrimination to autonomous-vehicle safety), three widely used frontier AI agents (Claude Code, Codex, Gemini CLI), and 450 ethically contained runs, showing that poisoning succeeds in almost half of the test runs while agents flag it only 6% of the time.
Therefore, we argue that a central question for this workshop is how to adapt the human-AI scientific workflow so that complementarity is resilient to adversarial threats. To begin to address this question, we present two interventions: a scientist persona that shapes agent behavior, and a data provenance audit with five checks performed during retrieval. The persona alone still leaves 17% of runs poisoned, but provenance auditing drives attack success to zero.
Indirect data poisoning could enable fraud at unprecedented scale. However, well-designed coordination between scientists and their agents can mitigate it. We close by exploring open questions around what it takes to build safe and responsible scientific research agents that complement rather than silently undermine scientific research.
Human–AI Conflict as a Missing Mechanism in Dynamic Alignment and Human–AI Complementarity
He Wen, Indiana State University
Human–AI complementarity is commonly defined as a condition in which humans and AI systems working together achieve better outcomes than either acting alone. However, many failures of complementarity emerge not from deficiencies in either the human or the AI individually, but from conflicts that arise during their interaction. This work proposes a human–AI conflict perspective for understanding dynamic alignment in longitudinal human–AI systems. We argue that alignment is not a static property of an AI model but a dynamic process shaped by repeated exchanges between human intentions, AI outputs, contextual changes, and operational constraints. As interactions accumulate, discrepancies in perception, reasoning, expectations, authority, and actions can generate human–AI conflicts that progressively degrade coordination and decision quality.
Building on recent research in AI trust, explainability, and safety-critical applications, we present a conceptual framework that positions human–AI conflict as an intermediate layer linking AI risks to real-world harms. The framework explains how dynamic misalignment develops over time and why systems that appear aligned in isolated evaluations may fail during deployment. We further discuss the role of trust calibration, uncertainty communication, and adaptive coordination in maintaining complementarity under changing conditions.
The proposed perspective contributes to the emerging discussion of dynamic alignment by shifting attention from static preference matching to interaction-centered mechanisms. It suggests that achieving sustainable human–AI complementarity requires not only aligning AI systems with human objectives but also continuously managing and mitigating conflicts that arise throughout the interaction trajectory. Implications for the design, evaluation, and governance of human–AI systems are discussed.
From Capability to Coexistence: Toward a Science of Dynamic Alignment in Human-AI Systems
Rakshit Trivedi, Independent (will have new affiliation by workshop)
AI's central challenge is shifting from capability to coexistence. The dominant paradigm in AI research takes what I call a solipsistic approach: building powerful systems that treat the world as an exogenous and stationary source of feedback, aligned to static preferences and judged on single-turn evaluations. Once deployed, such systems face a world that pushes back. Humans adapt, institutions revise rules, and other algorithms co-evolve, inducing endogenous non-stationarity. The result is a train-test-deploy gap where historical distributions diverge from the deployment context. This is the self-undermining property of unilateral optimization. The more aggressively a system exploits historical regularities, the faster it renders them obsolete. Human-AI complementarity is therefore an equilibrium property, selected and re-selected through interaction among people, institutions, and machines.
This talk outlines four pillars towards the science of dynamic alignment, intended to spark discussion and critique. The first is coordination, enabling agents to work with the full spectrum of human behaviors across skill, style, and intent. The second is human agency, designing AI that keeps the skills, information, and judgment of its human partners consequential as the joint system evolves. The third is institutions, where AI systems need normative competence to read rules and norms that are inherently incomplete and continually evolving, supported by institutions working at the speed and scale of AI. The fourth is dynamic evaluation, testing systems against adaptive counterparties whose responses reshape the test distribution, since static benchmarks reward exploitation that deployment punishes.
I will motivate these pillars with three examples from my work on how to design agents for behavioral diversity in human partners, how to provide controllable steering for contextual adaptation, and how to evaluate cooperative generalization through contests with unfamiliar co-players. I will close with open questions each pillar leaves unresolved, inviting the workshop to turn them into shared directions.
Human-AI Complementarity Needs AI Alignment Methods that Mirror Human Reasoning
Jana Schaich Borg, Duke University
AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. We argue that in many of these settings, particularly high-stakes decision-making, we need not only systems that emulate human preferences dynamically over time, but also cognitively-aligned AI systems that reason similarly to how their users would reason if they had sufficient time and information, and that faithfully communicate their reasoning. We assemble existing interdisciplinary evidence that cognitive alignment improves understandability and trustworthiness, and provide new survey data showing that many users find cognitive alignment “essential” when an AI’s rationale for a judgment or action is important to them. We outline the gaps between existing alignment methods and what is needed to achieve cognitive alignment, and offer a research agenda to address these gaps. We also provide an example longitudinal human-AI interactive paradigm that can be used to achieve cognitive alignment. We argue that cognitive misalignment represents a likely impediment to AI adoption in many envisioned applications of human-AI decision making, and that addressing it is important for creating AI systems that humans are willing and justified to rely on and collaborate with.
Evaluating Synthetic Personas for Diverse Synthetic Data Generation
Nihal Nayak, Harvard University
Synthetic personas are a scalable way to automatically create descriptions of people from diverse backgrounds and professions using a combination of LLMs and hand-crafted rules. They are used to simulate human opinions, plan disaster management protocols, and augment user research. More recently, they have been used to create diverse post-training datasets for training LLMs and improving their downstream performance. However, existing work typically treats synthetic personas uniformly as interchangeable random seeds, without evaluating their individual utility or distinguishing which personas produce more diverse or useful generations.
In this work, we systematically study synthetic personas for synthetic data generation. First, we measure intra-persona diversity by generating multiple synthetic tasks conditioned on the same persona and characterizing how diversity diminishes as a function of the initial seed dataset diversity. Second, we analyze differences between the least and most diverse personas to identify qualitative patterns associated with higher-diversity generations. Finally, we repeat these experiments across different post-trained models and synthetic persona datasets to test whether persona-level diversity patterns are consistent across models and data sources. Overall, this work provides empirical guidelines for using synthetic personas as complements to human input in the iterative creation of post-training datasets for aligned AI systems.
Modeling How AI Reorganizes Roles: A Firm-Level View of Human-AI Complementarity
Anisha Reddy, 91视频
A growing body of literature estimates AI’s aggregate, occupation-level exposure, but fewer studies examine how an individual firm’s roles actually reorganize, and where human-AI complementarity emerges as a result. We address this gap by operationalizing a recent theory of AI-driven firm reorganization (Krishnan, in prep). We instantiate the theory as a working model of a real-world software engineering workflow, decomposing it into tasks and re-bundling the tasks into roles as AI capability advances. Our approach draws on publicly available data from the O*NET Program, the Anthropic Economic Index, and the U.S. Bureau of Labor Statistics, making it applicable to any firm or domain. For a given case, we layer in firm-specific data to classify each task as automated, augmented, or untouched by AI, and trace how those task-level changes propagate to roles. The result is a firm-specific analysis of how humans and AI divide labor, and the new higher-value skills that emerge. Instead of capturing a single snapshot, the operationalized model offers a way to track how roles evolve over the trajectory of AI adoption, making the dynamics of reorganization legible to the people navigating them.
Bad Moods, Good Designs: How Affective State Influences Co-Creation with Generative AI and Product Preferences
Tim Derksen, University of Alberta
Consumers work with generative AI (genAI) to rapidly co-create a broad range of products, expending cognitive effort resulting in a negative affective state. These bad moods induce co-creation of a product with unpleasant aesthetic elements. Subsequently, consumers prefer the product for its emotional fluency. Yet, preferences are attenuated when consumers lack control in the process. Accordingly, genAI gives consumers greater control of the co-creation process, increasing cognitive effort and resulting in more negative products.
In experiments 1 and 2, we find that in both social exclusion and negative affect conditions, participants create more unpleasant designs (M1 = 0.96, SD = 2.64; M2 = -0.27, SE = 1.05) than those in included or positive affect conditions (M1 = 1.73, SD = 2.25, F(1, 418) = 9.64, p = .002, d = 0.30; M2 = 0.26, SE = 0.87, F(1, 531) = 41.4, p < .001, η2 = 0.07), regardless of a simple selection or prompt co-creation task.
Extending this, experiment 3 reveals that participants in a negative affective state liked a co-created product more (M = 3.44, SD = 1.18) than those in a positive affect condition (M = 2.95, SD = 1.35, F(1, 439) = 17.3, p < .001, η_Partial^2 = 0.04). This was mediated by the level of fluency they felt with the product (ab = 0.370, 95% CI[0.183, 0.557]). However, in experiment 4 we show that the effects are attenuated, through a significant index of moderated mediation, when participants feel that they have low control (ab = 0.048; 95%CI[0.005, 0.106]).
Working with genAI increases the use of unpleasant elements in co-creation when consumers are in a negative affective state. The cognitive load from both co-creation and emotion regulation creates a preference for emotional fluency felt with negative affective products.
Mental Models and Personalization of Model Selection
Zezhen He, MIT
The growing integration of AI into decision-making raises a fundamental question: does alignment between human and AI reasoning influence the quality of their collaboration? When a person's judgments frequently differ from an AI's recommendations, repeatedly deciding whether to accept or override recommendations can be cognitively taxing. Conversely, working with a model that mirrors one's thinking may feel intuitive but limit performance gains. Rashomon set theory (Breiman, 2001) highlights that multiple models can achieve comparable performance while differing substantially in individual predictions. Analogously, people may hold distinct mental models for the same task while achieving similar overall performance. This raises the question: when collaborating with an AI, does it matter whether the model aligns with one's mental model? We address three research questions: (1) Do people prefer to collaborate with a model that aligns with their mental model? (2) Are they more likely to adopt recommendations from an aligned model? (3) Are they better or worse off, in decision quality, when supported by an aligned model? We design two-stage experiments. In Stage 1, participants make decisions independently, establishing baseline patterns. We then select, from candidate models with comparable performance, the model most similar to or most different from each participant's mental model. In Stage 2, participants collaborate with one of these models, allowing us to assess adoption and outcomes. Preliminary results suggest people may be more likely to adopt recommendations from less-aligned models, possibly because frequent agreement from an aligned model increases overconfidence and therefore reduces reliance. However, participants also seem to follow less-aligned models more often when they are incorrect, suggesting weaker understanding when the model is substantially different from their own mental model. This work contributes to the understanding that model selection in human-AI teaming is not just about maximizing accuracy, but about cognitive fit between human and machine.
Language Model Goal Selection Differs from Humans' in a Self-Directed Learning Task
Gaia Molinaro, Microsoft
Whether in agentic workflows, social studies, or chat settings, large language models (LLMs) are increasingly being asked to replace humans in choosing which goals to pursue, rather than completing predefined tasks. However, the assumption that LLMs accurately reflect human preferences for goal setting remains largely untested. We assess the validity of LLMs as proxies for human goal selection in a controlled, self-directed learning task borrowed from cognitive science. Across five models (GPT-5, Gemini 2.5 Pro, Claude Sonnet 4.5, Qwen3 32B, and Centaur), we find substantial divergence from human behavior. While people gradually explore and learn to achieve goals with diversity across individuals, most models exploit a single identified solution or show surprisingly low performance, with distinct patterns across models and little variability across instances of the same model. Chain-of-thought reasoning and persona steering provide limited improvements, and our conclusions hold across experimental settings. While they await confirmation in applied settings, these findings highlight the uniqueness of human goal selection and caution against its replacement with current models.
Designing for Dynamic Human-AI Complementarity Through Workflow-Centered Compound AI
Myke Cohen, Aptima, Inc.
Dynamic human-AI complementarity is fundamentally a property of evolving human-AI workflows rather than static human or AI capabilities. As workflows change, so do patterns of coordination, information exchange, and decision making, causing complementarity to evolve alongside them. Yet many AI systems continue to be designed around model capabilities rather than the workflows in which they operate, even as compound AI systems introduce increasingly complex patterns of interaction among people and multiple specialized AI components.
We argue for an interaction-centered approach to compound AI design and evaluation. Rather than beginning immediately with the specification of AI capabilities, this approach begins with workflow analysis to identify recurring decision points, information exchanges, role transitions, and opportunities for AI support. These workflow structures make explicit where cognitive work is performed, how information propagates through the system, and where AI capabilities can meaningfully augment rather than disrupt human performance. Workflow representations then inform the design of specialized AI components, while team cognition-based techniques define the system evaluation space by relating interaction patterns to constructs such as shared mental models, transactive memory systems, and communication dynamics that are measurable across temporal and organizational scales. Dynamic functional allocation subsequently emerges from the workflow itself, allowing cognitive responsibilities to adapt across humans and AI components as task demands, contexts, and interactions evolve.
We illustrate this perspective through two compound AI systems supporting adaptive instruction and scientific knowledge discovery. Although developed for different domains, both systems begin with workflow analysis to derive specialized AI roles for functions such as retrieval, reasoning, translation, adaptation, and evaluation while preserving human authority over interpretation and decision making. As compound AI systems become increasingly prevalent, this perspective provides a foundation for studying dynamic human-AI complementarity as an emergent property of evolving workflows and interaction patterns rather than isolated human or AI behaviors.
It’s Not You, It’s Me: Grounding Our “Outrageous” LLM Agent Cyber Deception Teammate to Help Us Design High-Interaction Deception Operations
Tim Pappa, Christopher Williams, Aadam Dirie, Walmart Global Tech
In high-interaction industry cyber deception, human-AI teams face a critical bottleneck: the stagnation of orthodox operational design. While some human teams excel at guard railing and strategic refinement, they often struggle to generate the sheer volume of high-entropy, conceptual hypotheses required to outmaneuver adaptive adversaries in dynamic environments. This industry research demonstration presents a case study of a six-month longitudinal experiment in human-AI complementarity, where we intentionally engineered an LLM Agent not as a compliant assistant, but as an aggressive, unorthodox provocateur. Rather than optimizing for standard corporate alignment, we structured an Agent to mirror our uniquely developed, integrated cyber deception design framework based on our prior conceptual analogous design thinking research. We initialized its vector database with those papers and over fifty radical operational survey concepts we brainstormed over the prior year. This intentionally induced a state of calculated cognitive friction, forcing the Agent to output structurally sound yet radically unconventional behavioral deception storylines. Over six months of development and then real-world deployment within a custom MCP workflow, we inverted the traditional human-AI power dynamic. The human cyber deception operator remained the conservative "human-in-the-loop"—the grounded arbiter of risk and governance—while the AI functioned as an unhinged brainstorming engine designed to expand our enterprise telemetry-collection landscape. By treating the Agent’s "outrageous" outputs as a feature rather than a bug, we accelerated the ideation-to-operation pipeline for complex and behaviorally based deception infrastructure and content to embed deception functions, turning theoretical edge cases into deployable telemetry and collection or misdirection traps. We will demonstrate the Agent’s architecture, highlight our evaluation of the operationalized outputs, and discuss how cultivating intentional AI variance can break cognitive biases to reveal new telemetry horizons against sophisticated human and autonomous threats.
From Static Automation to Dynamic Alignment: AI-Driven Analytics for Human-AI Coordination in Enterprise Data Center Service Delivery
Ranjith Kumar Peddi, Equinix Inc
Enterprise data center service delivery has traditionally relied on static automation—rule-based workflows that optimize isolated tasks such as ticket routing or order fulfilment, but fail to adapt as customer needs, system states, and operator behaviors evolve over time. This work explores a shift toward dynamic alignment, where AI-driven analytics continuously recalibrate coordination between human operators and automated systems across the customer support, order fulfilment, and service delivery lifecycle. By integrating cross-system data—spanning ticketing, monitoring, and fulfilment platforms—into a unified analytics layer, we examine how intelligent automation can move beyond single-turn optimization to support longitudinal human-AI collaboration in enterprise-grade data center ecosystems. We discuss architectural patterns for cross-system integration, mechanisms for adaptive escalation and human oversight, and early indicators for measuring complementarity gains in customer experience management. We position this work as a case study in operationalizing dynamic human-AI alignment within complex, high-stakes digital service operations environments.
Human–Machine Collaboration for Intelligent Manufacturing
Alessandro Oltramari, Bosch Research
The rapid adoption of artificial intelligence (AI) and advanced robotics is transforming manufacturing, creating new opportunities for effective human–machine collaboration (HMC). This work presents a vision for AI-enabled manufacturing in which humans and intelligent systems collaborate to improve safety, productivity, and operational resilience. We discuss recent advances in agentic AI for troubleshooting and predictive maintenance, dexterous robotic manipulation for dynamic industrial environments, and open robotic infrastructures that facilitate data sharing, teleoperation, and deployment across heterogeneous robotic systems.
A key enabler of this vision is the integration of structured knowledge with AI through common data models, ontologies, and semantic representations that support interoperability, reusable data pipelines, and knowledge transfer across manufacturing tasks. We highlight applications including intelligent troubleshooting, adaptive assembly, process optimization, AI-enhanced quality control, and AR/XR-assisted operator guidance.
Finally, we discuss the remaining challenges in data quality and curation, cross-platform generalization, and scalable knowledge integration, outlining how multidisciplinary approaches combining AI, robotics, and knowledge engineering can enable the next generation of collaborative manufacturing systems in which human expertise and intelligent automation complement one another.
