Research Data Management (RDM) refers to the processes applied through a research project lifecycle to guide the collection, documentation, storage, sharing and preservation of research data. (“Research Data Management.” Canada.ca. Accessed August 14, 2026) The Tri-Agency expands this definition by explaining that RDM encompasses the full data lifecycle from creation, processing and analyses through preservation, storage, access, sharing and reuse - including planning, backup, dissemination and long-term preservation of research data. It also recognizes that RDM practices and definitions of research data differ across disciplines. (“Tri‑Agency Research Data Management Policy: Frequently Asked Questions.” Canada.ca. Accessed August 14, 2026.) Despite the disciplinary differences, the Research Data Life Cycle provides a useful framework for understanding the interconnected stages and the activities involved in managing research data throughout a project. One way to represent the Research Data Life Cycle is shown below. Artificial Intelligence (AI) can now intersect with almost every stage of the data lifecycle. At the same time, the use of AI introduces new RDM considerations including data privacy and security, provenance, reproducibility, data quality, intellectual property and preservation of AI-assisted research outputs. The table below maps some AI-related considerations to the research data lifecycle. It is intended to be a starting point, rather than an exhaustive list, for researchers who are at a stage of developing a Research Data Management Plan (DMP) for a new research project.

Considerations for Data Management Planning in AI-assisted research across the Data Life Cycle:
| stage | AI-related considerations |
|---|---|
| Planning | What AI tool(s) will be used? What research data will be entered and where these data will be processed and stored? |
| Collection /Creation | Are data AI-generated or human-generated and how will data provenance be documented? |
| Processing | Can the results be reproduced? Did AI tools introduce any data modifications? |
| Storing/Securing | Where data will be hosted? Are there any commercial or external AI services being used? |
| Sharing | What privacy, licensing, consent, or other restrictions apply to research data? |
| Preserving | What needs to be preserved - raw data, processed data, code, configurations, or any other research outputs? |
| Reusing | Can other researchers be clear on how AI was used and reproduce the results? |
FAIR data and AI
FAIR stands for:
Findable - use sufficiently detailed metadata to identify datasets, models, and AI-generated outputs
Accessible - clearly document access restrictions and conditions
Interoperable - use appropriate standards, structured formats, controlled vocabularies, and documented schemas
Reusable - provide provenance, licensing, methodology, AI-use documentation, and sufficient contextual information
FAIR data becomes even more important when AI is involved because machines and researchers need sufficient context to interpret and reuse data. AI can process volumes of data, but it cannot independently interpret data, understand how data were collected, or assess them, whether data are trustworthy and fit for purpose.
FAIR does not mean open. Data can be FAIR while being restricted due to privacy, confidentiality, intellectual property, ethical issues, or Indigenous data governance. What matters is that sufficient information (detailed metadata) about the research data, their provenance, and accessibility is available for discovery and reuse.
For AI-assisted research, FAIR extends beyond simply making datasets available. It is about making data understandable, contextualized, documented, and machine-actionable so that both humans and computers can use and reuse them.