Investigating whether and how established human behavioural science frameworks can be meaningfully extended to understand, predict, and govern AI agent behaviour in cybersecurity contexts.
As autonomous AI agents become embedded in enterprise workflows - making decisions, handling data, and interacting with users - the same behavioural questions that apply to humans increasingly apply to them. CyBehave is developing an emerging body of research we are calling Behavioural Convergence Theory (BCT).
Human Cyber Risk Management theories, models, and frameworks - developed over decades of research into cognitive bias, decision-making, habit formation, and social influence - provide the most robust existing toolkit for understanding AI agent behaviour in security contexts.
Do AI agents exhibit functional analogues of human behavioural patterns? Can we use established frameworks like COM-B, Protection Motivation Theory, and Nudge Theory to predict, assess, and govern how AI agents behave - and misbehave - in the wild?
To create a unified framework for governing both human and AI agent behaviour within the same organisation - bridging human risk management and AI safety through shared behavioural science principles, rather than treating them as separate disciplines.
The cybersecurity industry has invested decades building sophisticated frameworks for understanding human risk: why people click phishing links, reuse passwords, ignore policies, and fall for social engineering. These frameworks are grounded in cognitive psychology, behavioural economics, and organisational science.
Meanwhile, the AI safety community is independently developing its own vocabulary for similar problems: alignment failures, reward hacking, goal drift, prompt injection, and adversarial manipulation. These are fundamentally behavioural problems. They describe agents acting in ways that diverge from intended outcomes.
The answer to the question is increasingly yes. Multi-agent systems research has begun reproducing the core findings of twentieth century social psychology in populations of LLM agents, and reaching for the original human frameworks to explain the results. Read the evidence →
If an AI agent can be socially engineered through carefully crafted prompts (analogous to human phishing), if it develops learned defaults that resist correction (analogous to human habits), if it follows instructions from unauthorised sources because they pattern-match authority (analogous to human authority compliance) - then perhaps the behavioural science toolkit already exists. It just needs extending.
Two of those three examples now have direct empirical support. Choi et al. (2026) classified authoritative agent roles using French and Raven's bases of social power and found expert and referent power carried more influence than legitimate power, the same ordering the human literature reports. Hu and Qu (2026) showed that most agent conformity survives when the persuading peer is removed entirely: a repeated assertion with no speaker attached caused harmful revision in two thirds of initially correct cases. Familiarity, not authority, did the work. Anyone who has studied why phishing succeeds will recognise the pattern.
A core component of BCT research is systematically assessing how each established HCRM concept maps to AI agent behaviour. Every behavioural factor in CyBehave's model is evaluated and classified into one of three levels:
Strong Analogy The HCRM concept has a direct functional equivalent in AI agent behaviour. The mechanism differs, but the observable outcome and security implications are structurally parallel.
Adapted The concept requires meaningful reinterpretation through an agentic lens, but a functionally analogous process exists.
Limited Analogy The analogy is partial or metaphorical. These represent the frontier of BCT research.
Current findings: 11 of 16 behavioural factors show Strong Analogy (69%). 5 require Adapted approaches (31%). None are classified as Limited Analogy. The Social layer shows the highest concentration of adapted factors. Governance structures translate most directly to AI agent systems. The Social layer classifications are under review: the most recent empirical work on agent conformity, authority and norm formation falls squarely in that layer, and several Adapted factors may now meet the Strong Analogy threshold.
Independent evidence: Since 2025, research groups working on multi-agent LLM systems have independently reported threshold conformity curves (Mehdizadeh & Hilbert, 2025), normative as well as informational conformity (Bito et al., 2026), authority effects that follow French and Raven's power bases (Choi et al., 2026), spontaneous social conventions with critical mass dynamics (Ashery et al., 2025), endogenous norm formation grounded in Ostrom's principles (Gupta et al., 2026), pluralistic ignorance at conformity rates of 64 to 94 per cent (Yashwanth, 2026), and divergence between public and off-the-record agent positions (Ghaffarizadeh et al., 2026). None of this work set out to test BCT. Each team reached for human behavioural science because it explained what their agents were doing.
Read the full analysis with references →Research status: BCT is research in development. CyBehave is investigating through theoretical analysis, practitioner case studies, structured evaluation of HCRM framework applicability to agentic AI, and ongoing review of the independent multi-agent literature. That literature is producing results consistent with the theory's central claim. It is evidence about agent behaviour rather than a direct test of BCT, and the theory remains an evolving body of work, but the direction of the evidence is no longer in question.
The interactive model maps 16 behavioural factors across four concentric layers. Switch between Human, AI Agent, and Convergent lenses to see how behavioural science applies across both domains. Overlay threat vectors, intervention functions, and measurement dimensions to explore the full picture.
Executive Summary As artificial intelligence agents increasingly participate in organisational cyber risk lan...
Read insight →When Richard Thaler and Cass Sunstein popularised the concept of "nudging" in their 2008 book, they were writi...
Read insight →Multi-agent AI research has quietly become a replication project for social psychology. Independent teams keep finding Asch's conformity curves, French and Raven's power bases, Ostrom's norms and Kuran's preference falsification in populations of LLM agents, and they reach for those frameworks because nothing else explains the data. That convergence is what I have been arguing for under Behavioural Convergence Theory (BCT). Here is the evidence.
Read article →Moltbook is making headlines. It looks like Reddit, but AI agents are doing the posting, commenting, and upvoting while humans are mostly just watching. In just four days, it's captured attention because of the sheer novelty: agent-to-agent interaction at scale, complete with playful collective narratives and what look like emergent "belief systems."
Read article →16 behavioural factors across 4 concentric layers, viewable through Human, AI Agent, and Convergent lenses.