Citation verifiability
Page-level citations connect claims to their sources. Evidence previews and highlighted passages let researchers inspect the underlying material in context.
Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research
What does trustworthy AI look like in everyday evidence work? A five-month field study examines how policy and development professionals use an AI system that makes its sources and its limits visible.


Claims traceable to
specific passages.
An explanation of limits
and a constructive next step.
AVA brings a curated institutional corpus into the research workflow. Its approach to epistemic humility combines two complementary mechanisms.
Page-level citations connect claims to their sources. Evidence previews and highlighted passages let researchers inspect the underlying material in context.
When the available corpus cannot substantiate a response, AVA explains the evidence gap and suggests ways to refine the question or continue the search.

The five-month deployment involved over 2,200 policy and development professionals across 116 countries. The study combined platform logs, surveys, and 20 in-depth interviews to examine use in professional workflows. These figures describe different parts of the evidence: the deployment population is not the sample for every analysis.
Returning users reported greater time savings in an analysis by usage intensity. The randomized-invitation analysis found no statistically significant average impact on the measured productivity outcomes. §5.3 ↗
Platform logs, surveys, and interviews reveal how the system became part of professional work and where friction remained.
Participants used AVA for citable policy evidence, alongside general-purpose AI for tasks such as brainstorming. Use was shaped by the strengths and limits of each tool.
Workflow findings, §6.1Seventeen of the 20 interviewees rated citation verification 5 out of 5. Checking was selective, often triggered by surprising or numerical claims. High perceived usefulness does not establish that every response was correct.
Trust findings, §6.3Many participants read abstention as a sign of caution. Others experienced repeated non-answers as a barrier and moved to other tools. Explanations and query refinement mattered.
Abstention findings, §6.2The paper draws four lessons for specialised generative AI in evidence-intensive work.
Read the discussionSource quality, model behaviour, and interface design each contribute to trust. Keep residual error visible and verification easy to perform.
Use unanswered questions to identify gaps in the curated corpus. Interpret abstention alongside task fit, query reformulation, and abandonment.
Reduce the effort of inspecting claims and source context. Disclosure of AI use has a different role from checking the accuracy of an output.
The paper proposes handoffs that explain a system’s boundaries, preserve the user’s choice, and help them formulate a useful next query.
The five-month study does not establish long-term downstream impact. Survey and interview participation was voluntary, and predominantly English use limits conclusions about multilingual performance. §8 ↗
Example from Appendix I.3.
A guide to the paper’s relevance and evidentiary scope for researchers.
Nimisha Karnatak, Mohamad Chatila, Daniel Alejandro Pinzón Hernández, Reza Yazdanfar, Michelle Dugas, and Renos Vakis. 2026. Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26). Association for Computing Machinery. doi:10.1145/3772318.3791062.
Publication details and copyable BibTeX ↓Open-access paper ↗
Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research
@inproceedings{karnatak2026ava,
author = {Karnatak, Nimisha and Chatila, Mohamad and
Pinzón Hernández, Daniel Alejandro and Yazdanfar, Reza and
Dugas, Michelle and Vakis, Renos},
title = {Learning from AVA: Early Lessons from a Curated and
Trustworthy Generative AI for Policy and Development Research},
booktitle = {Proceedings of the 2026 CHI Conference on
Human Factors in Computing Systems},
year = {2026},
publisher = {Association for Computing Machinery},
doi = {10.1145/3772318.3791062},
url = {https://doi.org/10.1145/3772318.3791062}
}