Urban Wire How Can AI Responsibly Open Access to Government Data? An Evaluation of the National Secure Data Service Demonstration
Jameson Carter, Claire McKay Bowen, Aaron R. Williams, Rebecca John, Lindsay Ozburn, Liz Woolcott
Display Date

People working at computers with coding screens.

In 2022, the Advisory Committee on Data for Evidence Building recommended the federal government create a National Secure Data Service (NSDS) that would centralize information about the many federal datasets. Currently, these sources are scattered and siloed, which can obscure which data are available, what they can be used for, and how they should be characterized.

A future NSDS would provide a "government-wide set of shared services for data and evidence building,” including an online data concierge service that would provide a one-stop clearinghouse where a user can get information and connect to data across state and federal sources. To inform the creation of an NSDS, the National Science Foundation National Center for Science and Engineering Statistics administered the NSDS Demonstration projects, which funded the Urban Institute to evaluate whether a chatbot augmented with agentic AI can provide state and federal data to researchers directly as part of a data concierge service.

For the evaluation, we examined the NSDS data concierge service development process across several rounds of user testing within our team, external researchers, and state stakeholders. We primarily evaluated a chatbot augmented with agentic AI technology that would offer users access to public state and federal statistical data. From that user testing we developed a set of four principles that establish how the federal government could ensure chatbots for statistical data are accurate, promote best practices, and credit sources.

Four principles for the responsible development of an AI-supported data concierge

At a time when AI-augmented chatbots are proliferating, the lack of concrete federal standards has made responsible development difficult and uncertain. Urban’s previous work on AI informed our team's ability and expertise to undertake this evaluation. 

Recommendations discussed here have either been implemented, are underway, or should be considered in future efforts. For example, the Interagency Council on Statistical Policy notably made critical governance decisions that aligned with our recommendations, such as setting out testing requirements, identifying stakeholders, and defining implementation standards that the chatbot must meet.

The Home Page of the NSDS Chatbot

Navigator homepage with a chat-based government data search tool and a circular blue-and-red data visualization.

Source: National Secure Data Service, accesssed October 3, 2026. 

 

For the future development of NSDS data concierge services, we recommend developers adhere to the following principles:

1. Work with existing statistical systems

Data concierge services should augment current sources of quality data rather than replace them. Chatbots should credit and encourage users to review materials from official sources using structured resources (such as model context protocols) that help ensure data are correctly provided and contextualized to users.  

Developers of a data concierge service should

  • codevelop agency-specific model context protocols with data stewards to ensure transparency around the data used to inform replies and accuracy of chatbot interpretations of data requests (these protocols are technical infrastructures that help AI-augmented chatbots refer to predefined collections of data and metadata;
  • address the needs of users by identifying user personas and drawing on the expertise of data curators and other experts through user testing;
  • develop processes to triage requests so the chatbot knows when to answer a question directly and when to forward users to the NSDS or data curator staff;
  • ensure transparent, traceable replies through clear citations that help users understand a chatbot’s interpretations of data and the source data provided during a session; and
  • use deterministic tabulation code based on statistical systems’ best practices and configure chatbots to run that code. Otherwise, a chatbot may try to parse the data by generating code live, which introduces uncertainty.

2. Cultivate productive partnerships with stakeholders

Chatbots for the data concierge service should make it easier for stakeholders to access data, so actively testing and improving user experiences is crucial to successful development. Developers should create mutually beneficial partnerships with stakeholders in the data ecosystem, such as partners who steward key data sources, officials who access the data, and public policy researchers who analyze the data. As the chatbots pull from various sources, it is especially important to collaborate with data stewards to ensure accurate interpretation of those data and cultivate buy-in.

Developers of a data concierge service should

  • identify stakeholders that reflect the range of intended users and stewards of key data sources that will be used;
  • define early on what chatbots should and should not do to serve stakeholders—hard constraints should be differentiated from features that stakeholders inform;
  • co-design engagement, feedback, and decision-making processes with stakeholders;
  • clearly convey the chatbot’s comparative advantage in the current landscape of data access tools; and
  • respond directly and iteratively to feedback on a predefined basis during the development process.

3. Integrate quality evaluation at every step

Developers should evaluate how chatbots perform on prompt-response pairs of relevant questions and conduct qualitative user testing. Chatbots should be continually evaluated on a defined maintenance life cycle, using a mixture of quantitative and qualitative methods such that evaluators have a mechanism to report their results to government clients.

Developers of a data concierge service should

  • develop and adhere to a defined evaluation life cycle and apply best practices for measuring and responding to performance gaps during and after development;
  • devote resources for iterative evaluations among testers who reflect the range of intended end users and data stewards, so the chatbot remains accurate and useful;
  • establish clear, open feedback pathways for testers and integrate that feedback into the technology;
  • provide a systematic test plan for testers to ensure that key elements of technology are examined appropriately; and
  • define whether a chatbot should link data, how well critical context is linked to replies, and the conditions under which linking is acceptable.

4. Account for introduced privacy and security threats

Any repository of state and federal data must be able to withstand hacks and attempts to reveal personally identifiable information. During development, chatbot builders should stress test the security of any chatbot through red teaming. Data concierge service resources and data generated by the chatbot should be effectively firewalled from sensitive users and organizational information to reduce security risks.

Developers of a data concierge service should

  • design chatbots with privacy and security in mind early in the process—for example, a chatbot should not be allowed to help a user reconstruct a Census Bureau database;
  • use public data and ensure that resources organizing those data via model context protocols are publicly available;
  • design around privacy laws and regulations to ensure that the chatbot’s treatment of user data and chat logs complies with the relevant legal environment;
  • control model behavior to prevent chatbots from responding to sensitive questions and protect against prompt injections and other risks;
  • continually evaluate emergent risks introduced by chatbots powered with agentic AI and their supporting architecture; and
  • allow users to control how their personal data are treated and clarify the trade-offs between privacy and enhanced services.

What comes next?

As the NSDS Demonstration projects continue over the coming months, our team will publish lessons learned during the development process. Augmenting a chatbot with generative AI and aggregating disparate datasets introduces significant complexity, and we hope these lessons will contribute to a more responsible, effective path forward.

Research and Evidence Technology and Data Artificial Intelligence
Tags Bias and fairness in research, data, and technology Data collection Data governance and privacy Data resilience Evidence-based policy capacity Federal evaluation forum