The operational checks research teams should run before sharing digitized traditional-knowledge datasets
Digitizing traditional medicine knowledge for research or model training creates a workflow problem most institutions notice too late: the dataset moves through teams, vendors, and repositories, and somewhere along the way the record of who consented to what gets detached. By the time an AI team or a secondary researcher asks “can we use this for commercial training?” no one can answer with certainty.
Recent World Health Organization discussion on intellectual property, benefit sharing, and data governance for traditional medicine is a reminder that these questions are not abstract. They show up in the handoff between a research coordinator and a data vendor, or when an ethics review board asks whether contributors were told their knowledge might be licensed. The gap is not usually in the policy. It is in the per-person workflow that captures consent, records language preference, and keeps provenance metadata attached to the data as it moves.
The per-person flow that determines what you can do later
Think of one contributor. The workflow starts before anyone records anything. It includes how the person was invited, how consent was captured, what the consent language actually said, how the contributor’s preferred language was noted, and whether a structured record of those permissions travels with the audio file or transcript when it leaves your institution.
When that chain breaks, the dataset becomes fragile. Audio snippets collected in an Indigenous language are not just another row in a spreadsheet. They are tied to cultural protocols about who may share which knowledge and under what conditions. If the dataset only stores an English translation and the original verbatim is gone, the community cannot verify how their knowledge was represented. If the consent version is not recorded, a downstream user may assume broader rights than contributors agreed to.
Three items need to be captured at intake and stored as structured fields alongside the recording or transcript:
- Consent category and version, with a timestamp and the exact text the contributor saw.
- Language code and the original verbatim, plus any translation artifact and who produced it.
- Provenance metadata that answers which uses are permitted and which are not.
These are not fancy features. They are the minimum structured data an ethics review team or a downstream partner needs to filter the dataset by permitted use.
Three failure modes that repeat across projects
First, consent language is too generic. “Use for research” is not the same as “used for commercial model training” or “shared with third-party AI vendors.” The fix is to capture the contributor’s choice in clear categories at the time of consent so later requests for reuse can be filtered automatically. If your intake form does not ask “may we share this with commercial partners?” you cannot answer that question six months later when a licensing discussion starts.
Second, language and translation are treated as afterthoughts. If a contributor answers in their own language but the dataset only stores an English translation, the original nuance is lost. Store both the original verbatim and the translation, and tag who produced the translation and whether it was automated or human-reviewed. That metadata lets a reviewer verify accuracy and gives the community a way to check how their knowledge was rendered.
Third, provenance metadata gets detached. Datasets move through teams, vendors, and repositories. If the permit that allows a given use is not part of the dataset, someone later in the chain may assume broader rights than contributors agreed to. The question to ask before any transfer: when a file leaves our systems, does it carry an auditable record that answers which uses are allowed?
Design choices that shape what downstream partners can ethically do
How you collect and store traditional-knowledge inputs changes what you can do later. If a study collects healer interviews or community health knowledge, consider whether you will need to filter the dataset by contributors who opted into commercial use. That requirement should shape the initial intake form. It also affects research recruitment, because per-participant timelines and consent versions diverge across enrollees.
Offering voice and text-message collection alongside in-person or web forms can increase participation and improve representativeness, especially among contributors who are older, have limited internet access, or prefer a phone conversation. Storing the original voice recording plus a transcript lets reviewers retain cultural nuance and supports later verification if a benefit-sharing question arises.
Language access should be part of the design from the start. When open-ended feedback arrives in multiple languages, transcription and translation workflows need to be part of the data pipeline so one ethics review team can read responses consistently. Automated survey flows that let the contributor choose their language and record verbatim responses reduce downstream confusion and preserve the contributor’s voice. For background on multilingual design decisions, see multilingual outreach.
The system should produce an auditable record that can answer whether a contributor consented to a specific downstream use. That does not mean listing every technical detail. It means capturing the consent version, date, and the field that marks permitted use so a reviewer months later can answer “did this contributor allow reuse for model training?” without guessing.
Teams preparing to share or license traditional-knowledge datasets should run three checks before any transfer: is consent explicit about downstream uses, does each record carry language and provenance metadata, and can the receiving party demonstrate how they will enforce any use restrictions? Those checks are what turn policy into practice.
Related coverage: Intellectual Property and Traditional Medicine: Rethinking Access, Benefit Sharing and Data Governance — World Health Organization

