Strengthening Transparency at Research Institutions
Transparency at a research institution means that the people responsible for research can show what was done, when, by whom, and on what evidence, and that outsiders can check it
Strengthening Transparency at Research Institutions: What Verifiable Records Can (and Can't) Do
Transparency at a research institution means that the people responsible for research can show what was done, when, by whom, and on what evidence, and that outsiders can check it. Policies, training, open data and clear reporting lines all contribute. So does something less discussed: records that nobody, including the institution itself, can quietly rewrite later.
This article walks through what transparency actually requires, where institutions tend to fall short, and how verifiable records such as cryptographic timestamps fit in. They help with one specific part of the problem. They don't fix the rest, and we'll say clearly where that line is.
What does transparency mean for a research institution?
It helps to split it into two things that often get blurred together.
The first is openness: sharing methods, data, and results so others can scrutinise and reuse them. This is what most people picture when they hear "open science."
The second is accountability: being able to reconstruct, after the fact, what happened during a project. Which version of the dataset fed the published analysis? Was the protocol changed after the data came in? Who held the raw files in March?
An institution can be strong on one and weak on the other. A group might publish everything on release, yet be unable to show what its data looked like six months before. Or it might keep immaculate internal files that no outsider ever sees.
The European Code of Conduct for Research Integrity, published by ALLEA, covers both. Its four core principles are reliability, honesty, respect and accountability, and its guidance reaches from study design through data management to publication and peer review. The 2023 revision takes account of changes in data management practices, GDPR, and recent developments in Open Science and research assessment. The European Commission recognises the Code as the reference document for research integrity in EU-funded projects.
The Code also puts responsibility on institutions, not only individuals. Earlier editions asked institutions to ensure access to data is as open as possible and as closed as necessary, and to follow FAIR principles for data management. That's a useful phrase to hold onto, because it names the tension institutions live with: transparency pulls toward sharing, while confidentiality, privacy and commercial constraints pull the other way.
Where institutions usually fall short?
Rarely through bad intent. More often through ordinary friction.
Records live in too many places.
Raw data sits on a postdoc's laptop, analysis scripts in a private repository, protocol amendments in an email thread. When someone leaves, part of the story leaves with them.
Records can be edited without trace.
Shared drives and lab databases usually let authorised users overwrite files. Most of the time that's fine, even necessary. But it means a file's modification date is a statement the system makes about itself, not independent evidence.
Retention is vague.
Funders and regulators set minimum retention periods, and these vary. In the US, for example, institutions under Public Health Service misconduct rules must keep the institutional record and all sequestered evidence securely for seven years after a proceeding is completed (42 CFR § 93.318). University compliance offices point out that, beyond that, sponsor contracts may require longer retention than the baseline regulations. Many labs don't know which clock applies to them.
Integrity questions arrive late.
Allegations and queries typically come months or years after the work. At that point, the question "what did the data look like then?" is hard to answer with confidence.
Practical steps that actually strengthen transparency
None of this is exotic. Institutions that do it well tend to do several of these at once.
1- Write down a records policy people can follow. Say what counts as a research record, where it must be stored, who is responsible, and how long it's kept. Short and specific beats long and comprehensive.
2- Assign data stewardship. Someone in each unit should be able to say where a project's data is. Name a person, not a committee.
3- Version everything that matters. Datasets, protocols, analysis code, consent documents. Use version control where possible, and keep a changelog for what can't be versioned.
4- Register plans before the work. Preregistration of hypotheses and analysis plans, or registered reports, lets outsiders see what was intended before results existed.
5- Make amendments visible. Protocol changes happen. Hiding them is the problem, not the change.
6- Keep an independent anchor for key records. This is where verifiable timestamps come in. More below.
7- Rehearse the retrieval. Pick a closed project at random and ask someone to reconstruct its data lineage. Where it breaks is where your policy is weakest.
Where verifiable timestamps fit?
Say a lab freezes a dataset before analysis begins. The team wants to be able to show later that the frozen version existed on that date and hasn't changed since.
A cryptographic hash does half of that job. A hash function such as SHA-256 turns a file into a short fixed-length fingerprint. Change a single byte in the file and the fingerprint changes completely. So if you keep the fingerprint, you can later test whether a file is identical to the original.
The other half is the date. A hash on its own says nothing about when it was calculated. To make the timing independently checkable, the hash needs to be recorded somewhere nobody in the institution controls and nobody can retroactively edit. Public blockchains are one option, and trusted timestamping authorities using standards like RFC 3161 are another. Each has different trade-offs in cost, trust model and legal recognition.
How Certelo handles it?
Certelo is built around this pattern. The file stays on your device. The SHA-256 hash is calculated locally, and only that hash is anchored to the Electra Protocol blockchain. The original file isn't uploaded just to create the record.
For research institutions this matters more than it might in other settings. Unpublished data, patient-derived datasets, and pre-patent material are exactly the things you don't want to hand over to a third-party service just to get a date stamped on them. Keeping the file local removes that concern from the decision.
Later, anyone holding the original file can recompute its hash and check it against the blockchain record. If the hashes match, the file is identical to the one that was timestamped, and it existed no later than the recorded blockchain time.
A hypothetical example
This is illustrative, not a real case.
A clinical-research group finishes data collection for a study and exports the locked dataset as a single archive. Before anyone opens it for analysis, the data manager creates a Certelo record for the archive and saves the certificate information alongside the study's regulatory file. Eighteen months later, a reviewer asks whether the analysed dataset matches what was locked. The group recomputes the archive's hash and checks it against the blockchain record. A match supports their account. A mismatch would tell them, early and unambiguously, that something differs and needs explaining.
The same approach works for signed protocol versions, ethics approval letters, preregistration documents, instrument calibration logs, and code releases.
Important limitations
This is the part people skip, and it's why timestamping is sometimes oversold in integrity discussions.
A timestamp doesn't show the data is true. It shows that a particular file existed in a particular form by a particular time. Fabricated data can be timestamped just as easily as genuine data.
It doesn't establish authorship or ownership. The record links a file fingerprint to a time, not to a person. Showing who created something takes other evidence, such as lab notebooks, repository histories, institutional logs, or digital signatures.
It doesn't prove originality. Someone could timestamp a copy of another person's work. What the record can support is a claim about when, not whose.
It can't prove when something was created, only that it existed by then. The earliest provable point is the anchor time. The file may have existed far earlier.
Legal weight depends on jurisdiction and context. Whether a given court, regulator or funder accepts a blockchain timestamp as evidence varies, and this article isn't legal advice. Qualified electronic timestamps under EU rules, for instance, are a distinct legal category. Check with your counsel and with the specific body that will assess the record.
Process still matters. A hash taken before analysis is far more meaningful than one taken afterwards. The technology preserves whatever you feed it, so the discipline of when you do it comes from your institution's workflow.
Used inside a broader records policy, a timestamp is a small, sturdy piece of evidence. Used alone, it won't do much.
How to check a record?
Verification is meant to be possible without trusting the institution or Certelo's say-so.
1- Get the original file, or the hash recorded in the certificate.
2- If you have the file, recompute its SHA-256 hash locally.
3- Compare it with the hash in the certificate or blockchain record.
4- Check the blockchain record for the time at which the hash was anchored.
If the two hashes are identical, the file is bit-for-bit the same as the one that was timestamped. If they differ in any way, it isn't, though the check can't say what changed.
Frequently asked questions
Does timestamping research data count as research integrity compliance?
No. Compliance depends on the rules that apply to your funder, jurisdiction and field. Timestamping can support parts of a records practice, but it doesn't replace policies, training, ethics review or retention schedules.
Can a blockchain timestamp prove who created a dataset?
No. It records that a file fingerprint existed at a time, not who produced the file. Authorship evidence needs additional records.
Do we have to upload our data to timestamp it?
Not with Certelo. The hash is computed on your device, and the original file isn't uploaded to create the record. Other services work differently, so check before using them with sensitive data.
What happens if the file is changed after it's timestamped?
Its hash will no longer match the recorded one. That's the point: any change, even a trivial one, is detectable. Keep the exact original file, since the record can't recreate it.
Is this the same as preregistration?
No, they complement each other. Preregistration publicly states plans in a registry. A timestamp helps show that a specific file existed unchanged at a specific time. You might do both.
How long should an institution keep research records?
It depends on the funder, regulator, contract and country. The US rule cited above sets seven years for misconduct-proceeding records, and sponsors may require more. Your research compliance office is the right place to confirm the applicable period.
Can this help with disputes about prior work or priority?
It can contribute evidence about when a document existed. Whether that settles a dispute depends on the other evidence and on who is deciding.
A closing note
Institutional transparency is mostly unglamorous: clear policies, named people, versioned files, and the habit of writing things down at the time. Verifiable timestamps add something specific to that mix, which is an independent answer to "can you show this file hasn't changed?" For teams working with sensitive material, being able to get that answer without handing the file over is a practical advantage.
If your institution wants to try it on a real project, Certelo's verification page lets you see how a record is checked before you create one.