CHI ’26Human–computer interaction research

Learning from AVA

Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research

Introduction

What does trustworthy AI look like in everyday evidence work? A five-month field study examines how policy and development professionals use an AI system that makes its sources and its limits visible.

AI + VERIFIED ANALYSIS01 / SYSTEM LOGIC
A research questionWhat does the evidence say?
4,000+ World Bank reportsA curated evidence base
Retrieve · synthesise · verifyIs there sufficient
supporting evidence?
Evidence supports an answerA source-linked
response

Claims traceable to
specific passages.

Evidence is insufficientA reasoned
abstention

An explanation of limits
and a constructive next step.

Figure 1. AVA answers from a curated corpus and makes its evidence boundaries explicit. Sufficient evidence leads to a source-linked response; insufficient evidence leads to reasoned abstention.Simplified from §3

System Architecture

AVA brings a curated institutional corpus into the research workflow. Its approach to epistemic humility combines two complementary mechanisms.

Citation verifiability

Page-level citations connect claims to their sources. Evidence previews and highlighted passages let researchers inspect the underlying material in context.

Reasoned abstention

When the available corpus cannot substantiate a response, AVA explains the evidence gap and suggests ways to refine the question or continue the search.

Figure 3 from the paper: AVA's interface connects a research query to inline citations, evidence previews and a source document viewer, with options to save and edit notes.
Figure 3 from the paper. From asking a question to checking sources and saving evidence. Karnatak et al., 2026. Source ↗ · CC BY 4.0

Methodology

The five-month deployment involved over 2,200 policy and development professionals across 116 countries. The study combined platform logs, surveys, and 20 in-depth interviews to examine use in professional workflows. These figures describe different parts of the evidence: the deployment population is not the sample for every analysis.

Study methods, §4 ↗

Quantitative Results

CORPUS COVERAGE AND ABSTENTIONView full-size figure ↗
Figure 5 plots five-day total query volume in blue against the percentage of queries receiving no response in red. Abstention is high during early deployment, drops sharply in July after corpus expansion, and remains comparatively low as query volume fluctuates.
Figure 5 from the paper. Abstention falls sharply after corpus expansion. The two lines use different axes: query counts on the left and the no-response percentage on the right. Karnatak et al., 2026 ↗

Impact Evaluation: Improvements to Work Efficiency (RQ3)

Returning users reported greater time savings in an analysis by usage intensity. The randomized-invitation analysis found no statistically significant average impact on the measured productivity outcomes. §5.3 ↗

Qualitative Results

Platform logs, surveys, and interviews reveal how the system became part of professional work and where friction remained.

Positioning AVA within Evidence-Based Workflows (RQ1)

Participants used AVA for citable policy evidence, alongside general-purpose AI for tasks such as brainstorming. Use was shaped by the strengths and limits of each tool.

Workflow findings, §6.1
17 / 20

Trust Calibration in Evidence Practices (RQ2)

Seventeen of the 20 interviewees rated citation verification 5 out of 5. Checking was selective, often triggered by surprising or numerical claims. High perceived usefulness does not establish that every response was correct.

Trust findings, §6.3

Reasoned Abstention: When AI Says ‘I Don’t Know’ (RQ2)

Many participants read abstention as a sign of caution. Others experienced repeated non-answers as a barrier and moved to other tools. Explanations and query refinement mattered.

Abstention findings, §6.2

Discussion

The paper draws four lessons for specialised generative AI in evidence-intensive work.

Read the discussion
  1. Lesson 1: Designing Specialized AI Systems for High-Stakes Knowledge Work Requires an End-to-End Trust Pipeline

    Source quality, model behaviour, and interface design each contribute to trust. Keep residual error visible and verification easy to perform.

  2. Lesson 2: Balancing the Source Quality and Coverage Tradeoff

    Use unanswered questions to identify gaps in the curated corpus. Interpret abstention alongside task fit, query reformulation, and abandonment.

  3. Lesson 3: Design for Verification, Not Just Disclosure

    Reduce the effort of inspecting claims and source context. Disclosure of AI use has a different role from checking the accuracy of an output.

  4. Lesson 4: Embracing the Future: From a Single Tool to a Collaborative AI Ecosystem

    The paper proposes handoffs that explain a system’s boundaries, preserve the user’s choice, and help them formulate a useful next query.

Limitations

The five-month study does not establish long-term downstream impact. Survey and interview participation was voluntary, and predominantly English use limits conclusions about multilingual performance. §8 ↗

Epistemic Humility: A Critical Divergence

Example from Appendix I.3.

OUT-OF-DOMAIN STRESS TESTView full-size figure ↗
Figure 6 shows Perplexity answering an out-of-domain request for a pizza recipe, with citations to policy documents. The paper reports that the policy corpus contains no supporting evidence for this response.
Figure 6 from the paper’s appendix. In the reported “pizza recipe” stress test, Perplexity supplies an answer despite a lack of supporting evidence in the policy corpus. This example illustrates why the presence of citations alone does not establish evidential support; it is not an estimate of overall error rates or current product performance. Karnatak et al., 2026 ↗

Research contributions and when to cite

A guide to the paper’s relevance and evidentiary scope for researchers.

What does this paper contribute?

  1. Large-scale, real-world HCI evidence. A five-month, multi-institutional field study of epistemic humility in a deployed generative AI system, involving over 2,200 policy and development professionals across 116 countries. Platform logs, surveys, and 20 interviews examine use in everyday professional workflows. Methods, §4.
  2. Evidence on humility and verification in practice. Findings on how users interpret reasoned abstention, selectively inspect page-anchored citations, and develop reliance on a source-bounded AI system. The study identifies both the value of explicit limits and the friction caused by unanswered questions. Findings, §6.
  3. Design lessons for specialised AI. An account linking corpus curation, abstention, and interface-level verification, alongside lessons on source quality and coverage. The paper proposes ecosystem-aware humility and intelligent handoffs as directions for collaboration between specialised and general-purpose tools. Discussion, §7.

When is this paper relevant to cite?

Large-scale and in-the-wild evaluation in HCI
Research on multi-month, multi-institutional deployment studies of generative AI with real-world users, particularly studies of epistemic humility in professional knowledge work.
Epistemic humility and reasoned abstention
Research on making a generative system’s knowledge boundaries explicit and understanding how users respond to unsupported queries, explanations of limits, and constructive redirection.
Citation verification and calibrated reliance
Research on page-level citations, source provenance, evidence inspection, and the role of verification interfaces in human–AI collaboration.
Specialised AI for policy and development
Research on curated institutional corpora, evidence-grounded synthesis, and the trade-off between source quality and the breadth of questions a system can address.
Multi-tool professional workflows
Research on how people combine specialist and general-purpose AI tools, and design proposals for responsible handoffs between them.

Canonical reference

Nimisha Karnatak, Mohamad Chatila, Daniel Alejandro Pinzón Hernández, Reza Yazdanfar, Michelle Dugas, and Renos Vakis. 2026. Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26). Association for Computing Machinery. doi:10.1145/3772318.3791062.

05 / THE PUBLICATIONCHI ’26 · Barcelona, Spain · 13–17 April 2026

Learning from AVA

Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research

Nimisha KarnatakUniversity of Oxford

Mohamad Chatila · Daniel Alejandro Pinzón Hernández
Michelle Dugas · Renos Vakis
The World Bank Group

Reza YazdanfarNouswise, Inc.

Cite this paper / BibTeX
@inproceedings{karnatak2026ava,
  author = {Karnatak, Nimisha and Chatila, Mohamad and
    Pinzón Hernández, Daniel Alejandro and Yazdanfar, Reza and
    Dugas, Michelle and Vakis, Renos},
  title = {Learning from AVA: Early Lessons from a Curated and
    Trustworthy Generative AI for Policy and Development Research},
  booktitle = {Proceedings of the 2026 CHI Conference on
    Human Factors in Computing Systems},
  year = {2026},
  publisher = {Association for Computing Machinery},
  doi = {10.1145/3772318.3791062},
  url = {https://doi.org/10.1145/3772318.3791062}
}