The successful training and deployment of artificial intelligence (AI) tools are highly dependent on the quality of their input data. As the proliferation of AI tools for healthcare applications continues, it is critically important that stakeholders consider the provenance and quality of any data used to power an AI application. Healthcare AI model output success depends on data selection, for small differences in initial inputs may produce vastly different outcomes. There are many different types and flavors of healthcare data, each suitable for some tasks and not for others, making data selection difficult. In this overview, we describe some of the considerations at play in bringing the right data to bear for a given use.
Our discussion points include the following.
- Data confidence: Selecting the right data
- Variance in healthcare claims data: Examining open, closed, and paid claims
- Reference data: Providing additional context by reviewing entities, geographies, or concepts involved in claims and enrollment data