Healthcare AI Data Provenance: Seven Questions Before You Connect a New Data Feed

Author: Eunoia Consulting Co. | Published: October 7, 2026

Before connecting a new healthcare data feed to an AI workflow, ask seven provenance questions about source authority, transformations, quality, access, change control, fallbacks and the response required when a data condition changes.

Key Takeaways

  • A new data feed can change an AI workflow’s meaning, timeliness and risk profile.
  • Start with the decision the data will influence, not the integration alone.
  • Document source ownership, transformation, quality limits and permitted purpose.
  • Define triggers for reassessment when source systems or mappings change.
  • Use a fallback whenever provenance or quality cannot be confirmed.

A new data feed can make a healthcare AI initiative look ready: more fields, more history, more context, more apparent completeness. But a data connection is not merely a technical event. It can change the meaning, timeliness, quality, privacy posture, workflow role and risk profile of the system that receives it.

Data provenance is the ability to explain where a piece of information came from, how it was produced or changed, and which people, systems and processes influenced it. In the FHIR specification, a Provenance resource records entities and processes involved in producing, delivering or otherwise influencing a resource. That technical definition is useful, but a healthcare AI programme needs an operating version too: can an accountable team explain what data the system used, why it was available, how it was transformed and when its meaning or reliability might have changed?

This article presents seven practical questions to ask before connecting a new feed to a healthcare AI workflow. It is not a claim that provenance alone makes a system safe, compliant or appropriate. It is a decision discipline that helps organisations recognise when a data change needs technical validation, clinical input, privacy or security review, vendor clarification or a pause.

Why provenance matters before an AI data connection

A data feed can introduce changes that are invisible at the interface level. Field names may match while definitions differ. A record may arrive with a delay that changes its operational relevance. A transformation may remove a qualifier a reviewer needs. A new source may be complete for one site and sparse for another. A vendor might update a mapping, a code set or an extraction schedule without a clear signal to downstream users.

For AI-enabled tools, those changes can influence what an output represents. A summary may omit a newly unavailable data element. A queue-prioritisation feature may receive a population different from the one originally reviewed. A document assistant may draw context from a source that contains stale or duplicated content. The question is not only “Does the integration work?” It is “Does the integration preserve the operating assumptions on which the intended use and controls depend?”

The NIST AI Risk Management Framework is voluntary and intended to help organisations incorporate trustworthiness considerations into AI design, development, use and evaluation. Its focus on documenting context, assumptions, limitations and risk-management activity gives teams a useful way to treat a data connection as more than an IT ticket.

> Eunoia recommendation: No team should have to rely on a generic assurance that a feed is “live.” For a material AI workflow, the accountable owner should be able to identify the source, permitted purpose, transformations, quality checks, known limitations and response path when the feed changes.

1. What decision or workflow will this data influence?

Start with the decision context—not the data set. Ask exactly where the information will appear, who will use it and what they may do differently because of it.

A feed that supports a non-clinical operational summary has a different consequence from a feed that influences a clinical review, patient-access prioritisation, documentation draft, revenue-cycle work queue or leadership dashboard. The same field can carry different risks depending on the workflow.

Document:

  • the AI-enabled system, feature and configuration in scope;
  • the intended use of the output;
  • the user role and whether a human review is required before action;
  • the setting, patient population, sites and service lines in scope;
  • the decisions or actions the output can influence; and
  • the uses that are explicitly out of scope.
  • The ASTP/ONC HTI-1 Final Rule includes algorithm-transparency requirements in the certified-health-IT context and describes a baseline of information intended to help clinical users assess fairness, appropriateness, validity, effectiveness and safety. It does not create a universal local approval checklist for every AI tool, but it reinforces the value of clear information about what an algorithm does and the context in which it is used.

    2. What is the authoritative source, and who owns it?

    A source name is not enough. Identify the system of record, the business or clinical owner, the technical owner and the person responsible for granting or revoking access.

    A provenance record should distinguish, for example:

  • primary clinical documentation from a derived report;
  • a scheduling or billing system from a manually maintained spreadsheet;
  • a source integration from a vendor-provided data export;
  • a production environment from a test or training environment; and
  • a source that is authoritative for a specific field from a source that merely reproduces it.
  • Where more than one source supplies a similar field, document the precedence rule. A tool may handle dates, medication lists, problem lists, appointments, locations or payer information differently when values conflict. If the team has not decided which source is authoritative, the connection is not ready to support an automatic interpretation.

    3. What transformations occur before the AI system receives the data?

    Raw data rarely arrives unchanged. It may be filtered, mapped, normalised, joined, de-identified, enriched, scored, deduplicated, truncated or delayed. Each step can change the meaning of the result.

    Create a short transformation record:

    | Stage | Questions to answer | |---|---| | Extraction | How is information selected, how often, and under what access authority? | | Transfer | Which integration, interface, batch process or file mechanism moves the information? | | Mapping | How are identifiers, code sets, field names and values translated? | | Processing | Is information filtered, summarised, joined, normalised, de-identified or otherwise changed? | | Delivery | When does the AI system receive it, and how can users recognise a delay or failure? | | Retention | Where is the copied or transformed information stored, for how long and under whose control? |

    The record does not have to expose sensitive content. It should be detailed enough for an accountable reviewer to understand which transformation might explain an unexpected output. If a vendor performs a material transformation, seek enough documentation to evaluate its operational effect rather than accepting a high-level statement that data is “processed securely.”

    4. What are the known completeness, freshness and quality limits?

    Data quality is not a single score. A field can be accurate for one site but not another, complete for active patients but not historical records, or timely in the source system but delayed in the downstream feed.

    For the initial connection, define a small validation set. It can include:

  • expected record volumes by source and site;
  • completeness of critical fields for the intended use;
  • observed time from source update to downstream availability;
  • duplicate, conflicting or unmapped records;
  • handling of null, unknown, retired or out-of-range values;
  • reconciliation samples against the source system; and
  • known data gaps that users must understand.
  • Do not convert limited validation into a broad accuracy claim. The purpose is to identify whether the feed is fit for a defined operational use and whether the limitations are visible to the people who review outputs.

    5. Does the connection preserve purpose, access and privacy boundaries?

    A technically possible connection may still be outside the organisation’s agreed purpose or access model. Before activating a feed, confirm the actual data elements, authorised users, vendor or internal roles, data locations, access controls, retention arrangements and third-party dependencies.

    This should be assessed through the organisation’s applicable privacy, security, contracting and governance processes. A data-governance article cannot determine legal obligations or provide a universal answer. It can, however, prompt practical questions:

  • Is each data element necessary for the stated use?
  • Is there a documented purpose for the connection and downstream processing?
  • Which party can access, download, change or retain the data?
  • What happens when a user, site or vendor relationship changes?
  • What incident, correction or deletion path applies if data is sent in error?
  • The NIST Generative AI Profile describes data provenance, data protection, data retention, incident response, monitoring and third-party considerations as governance tools organisations may apply to generative AI contexts. That is a useful prompt to connect data architecture with operating accountability.

    6. How will the team detect and respond to feed changes?

    A reliable first connection can become unreliable when a source system, interface, vendor, code set, data mapping or operating process changes. Define the triggers that require a reassessment before the feed continues to support the same AI workflow.

    Typical triggers include:

  • a vendor or source-system release;
  • a changed integration version, authentication method or data contract;
  • an altered mapping, field definition, value set or data dictionary;
  • a material reduction in data completeness or freshness;
  • a new site, service line, population or user role;
  • an incident, security concern, privacy concern or repeated user report; and
  • a request to use the feed for a new purpose.
  • Use a change-control record to capture the trigger, what was affected, evidence reviewed, decision owner, conditions and follow-up. Eunoia’s healthcare AI change-log framework provides a practical way to connect a changed data condition to a documented decision rather than treating it as background technical noise.

    7. What is the fallback when provenance or quality cannot be confirmed?

    A dependable workflow needs a defined response for uncertainty. If the team cannot confirm the source, freshness, transformation or completeness of data that materially informs the output, the safest action may be to withhold that output from the intended workflow until the issue is understood.

    Document the fallback in advance:

  • can the user return to the source system or existing process?
  • can the AI function be restricted to a lower-risk support role?
  • who decides to pause, bypass or restore the connection?
  • how are affected staff informed?
  • what evidence is required before the feed returns to use?
  • how will the incident or change be reviewed for a recurring pattern?
  • The fallback should be proportionate to the workflow. A non-clinical administrative summary and an output that supports a consequential decision may require different controls. What matters is that the response is explicit and owned before an unexpected issue occurs.

    A practical provenance packet for a new feed

    Before activation, assemble a concise packet containing:

  • intended use and out-of-scope uses;
  • source-system and business-owner details;
  • field inventory and minimum-necessary rationale;
  • transformation and mapping description;
  • data-quality validation approach and known limitations;
  • access, retention and third-party responsibility overview;
  • change triggers, monitoring indicators and escalation path; and
  • approval record, accountable owner and next review date.
  • This packet is not a substitute for formal architecture, security, privacy or clinical validation processes. It is the connective tissue that helps those processes work from the same facts.

    Make provenance usable by the people who need it

    Good provenance is not a hidden technical diagram. It is accessible enough for operational and clinical leaders to understand what an AI-enabled workflow is using and what to do when that information changes. It turns a data feed from an opaque dependency into a reviewable operating asset.

    Eunoia Consulting Co. helps healthcare organisations establish decision-ready data-governance practices for AI implementation, including source mapping, data-quality checks, vendor questions, change control and operational accountability. To discuss a structured approach for a defined healthcare workflow, explore Eunoia’s data-governance advisory services or contact the team. This is advisory support for internal decision readiness, not legal, clinical, privacy, security or regulatory advice.

    Sources

  • HL7 FHIR — Provenance
  • NIST — AI Risk Management Framework
  • NIST — Generative AI Profile
  • ASTP/ONC — HTI-1 Final Rule
This article was produced by the Eunoia Consulting Co. Editorial Team. Eunoia Consulting Co. specialises in AI governance, healthcare operations, and data strategy for healthcare and veterinary organisations.