How Human Clarity Institute Builds AI Behaviour Benchmarks

Human Clarity Institute develops AI Behaviour Benchmarks by combining original behavioural research with independent academic and population research. This page explains the methodology used across the Benchmark Library.

Every Human Clarity Institute benchmark combines three evidence streams: original HCI behavioural benchmark data, independent academic and population research, and behavioural synthesis to identify the strongest and most consistent patterns of human adaptation to artificial intelligence.

Why We Build Behaviour Benchmarks

Artificial intelligence is changing how people think, decide, learn, work and relate to the world around them. Many organisations measure AI adoption, usage volume or public attitudes toward AI. Fewer measure how human behaviour changes as AI becomes part of everyday life.

Human Clarity Institute benchmarks are designed to provide evidence-based reference points for understanding those behavioural changes. Each benchmark brings together multiple forms of evidence to identify patterns in human adaptation to AI, including trust, reliance, verification, decision-making, agency, thought partnership and emotional use.

Three Evidence Streams

Human Clarity Institute benchmarks do not rely on a single dataset or study. They are built by integrating original HCI behavioural data with independent research to identify the strongest and most consistent patterns of human adaptation to artificial intelligence.

Original HCI Behavioural Research Independent Academic and Population Research Behavioural Synthesis Evidence-Based Benchmark Layer 1 Layer 2 Layer 3 Output

Each benchmark combines original HCI data, independent research and behavioural synthesis before being published as an evidence-based benchmark.

Original HCI Behavioural Research

Original benchmark datasets measure how people interact with artificial intelligence across core dimensions of human behaviour, including trust, reliance, verification, decision-making, agency, emotional regulation, thought partnership, disclosure and relationship with AI.

HCI benchmark research currently draws on 10,500+ participants across 21 studies, using Prolific research panels and samples across the United States, United Kingdom, Canada, Australia and Ireland. Sample composition, study design and interpretation vary by benchmark topic and are described where relevant on individual benchmark pages.

How Research Data Supports HCI Assessment Products

Anonymised research responses from earlier HCI studies are also used internally to establish benchmark distributions for HCI assessment products, including paid products. These distributions allow an assessment participant’s reported behaviour to be compared with patterns observed across HCI research samples.

Historical research participants are not identified in HCI assessment products. Participant-level research records are not provided to customers, and assessment outputs present only aggregated comparisons, such as response distributions, percentile positions and cohort-level reference points. The underlying benchmark data is maintained internally and is not made available through the assessment product.

This use is disclosed so the relationship between HCI’s research programme and its commercial assessment products is clear.

AI assessment reference: For official assessment terminology, dimension definitions and interpretation limits, see the HCI AI Identity & Behaviour Assessment AI Reference .

Independent Academic and Population Research

Benchmarks are compared against independent evidence from peer-reviewed studies, university research, population surveys, government research, and trusted research organisations. Examples include sources such as Pew Research Center, Brookings Institution, OpenAI, DataReportal and academic research on human-AI interaction.

Behavioural Synthesis

Behavioural synthesis compares HCI benchmark findings with independent academic and population research to identify where evidence consistently converges, where it differs, and where important new behavioural patterns are emerging.

Benchmark Development Process

Every benchmark begins with a clearly defined behavioural question. Evidence is then gathered from HCI benchmark datasets, independent academic research and population studies before being synthesised into a single evidence-based benchmark.

BehaviouralQuestion EvidenceCollection BehaviouralAnalysis BenchmarkDevelopment Ongoing Review

Benchmarks are developed from behavioural questions, evidence collection, analysis, synthesis and ongoing review.

Published benchmarks are reviewed as new evidence emerges. Updates may reflect new HCI datasets, new academic findings, new population research or new Human Signals Observatory signals.

Principles Behind Every Benchmark

Independent

Findings are reported without commercial influence over the interpretation of results.

Evidence-based

Benchmark conclusions are supported by multiple evidence streams rather than a single study or isolated statistic.

Behaviour-focused

Benchmarks measure how people change in relation to AI. They are focused on human behaviour, not AI capability.

Conservative

Claims are written to reflect what the evidence shows. Where findings are uncertain, emerging or limited, that uncertainty is acknowledged.

Transparent

Benchmark pages cite source material, link to relevant research where available and explain how findings should be interpreted.

Continuously reviewed

AI systems and human behaviours change over time. Benchmarks are reviewed and updated as new evidence becomes available.

Evidence Sources

Every benchmark draws on multiple evidence types to build a complete picture. HCI combines original behavioural research with independent academic findings, population surveys and trusted research institutions. This diversity of evidence strengthens conclusions and reveals where different sources agree, where they diverge and where further research is needed.

Evidence types may include:

  • Human Clarity Institute benchmark datasets
  • Academic journals and peer-reviewed research
  • University research programmes
  • Population surveys and polling organisations
  • Government and public-interest research bodies
  • Independent research organisations
  • Human Signals Observatory research scans
  • High-quality reporting where it identifies emerging human-AI behavioural signals

Keeping Benchmarks Current

AI is changing rapidly, and human behaviour changes with it. Benchmarks are living publications and may be reviewed or updated as new evidence becomes available from HCI research, academic literature, population studies and Human Signals Observatory scans.

Updates may include new statistics, revised interpretation, additional external sources, new related benchmarks or clearer explanations of limitations.

Interpreting Benchmark Findings

Benchmarks describe population-level behavioural patterns. They are not diagnoses, clinical assessments or deterministic predictions about any one individual. A benchmark can show how behaviour commonly varies across a population, but individual context still matters.

The purpose of a benchmark is to provide context. It helps readers understand whether a pattern appears common, uncommon, emerging or strongly associated with particular forms of AI use.

Limitations

Human behaviour varies between individuals and changes over time. AI technologies, usage patterns and social norms also evolve quickly. Benchmark findings should therefore be understood as evidence-based reference points rather than fixed rules.

Survey-based findings may be influenced by sample composition, wording, self-report limits and the time at which the data was collected. HCI research samples are recruited through Prolific and are not presented as nationally representative general-population norms unless an individual benchmark page expressly states otherwise. External research also varies in design, quality and generalisability. HCI benchmarks are developed to make these patterns clearer, while acknowledging uncertainty where it remains.

Related Resources

Page information

Published: July 2026

Last updated: August 2026

Applies to: Human Clarity Institute AI Behaviour Benchmarks