Hashing Documents Instead of Storing Them
You can record a document's SHA-256 hash instead of storing the document itself, and for many purposes that is enough. A hash is a short fixed-length fingerprint calculated from the file's exact bytes
Hashing Documents Instead of Storing Them: How It Works and Where It Falls Short
You can record a document's SHA-256 hash instead of storing the document itself, and for many purposes that is enough. A hash is a short fixed-length fingerprint calculated from the file's exact bytes. It lets you check later whether a file is identical to the one you hashed, and it doesn't reveal the file's contents. What it cannot do is replace the file: the hash can't be turned back into the document, so someone still has to keep the original.
That trade-off is the whole story, and the rest of this article walks through it.
What a document hash actually is?
A cryptographic hash function takes any input, whether a one-page letter or a 4 GB video, and produces a fixed-size output. SHA-256, specified by NIST in FIPS 180-4, always produces 256 bits, usually written as 64 hexadecimal characters.
Three properties make this useful for documents:
* The same file always produces the same hash.
* Changing the file in any way, even one character or one pixel, produces a completely different hash.
* Working backwards from the hash to the file is not computationally feasible, and finding a different file with the same hash is not either.
So a hash works like a fingerprint. You can't rebuild a person from a fingerprint, but you can confirm whether a given person matches it.
Why not just store the document?
Storing documents is simple, but it comes with costs that people often underestimate. Every copy is something to secure, back up, and eventually delete. Under the GDPR, the data minimisation principle in Article 5(1)(c) says personal data should be limited to what is necessary for the purpose. If your purpose is only to show later that a file hasn't changed, a hash can do that without you holding the content.
There is also a trust issue. If you hand a contract, a draft, or a research dataset to a third-party service just to get it timestamped, you've shared the contents with that service. A hash lets you share proof without sharing substance.
How it works in practice?
The workflow is short:
1. Calculate the SHA-256 hash of the file on your own device.
2. Record the hash somewhere that can be checked later, ideally a place nobody can quietly edit afterwards.
3. Keep the original file safe.
4. When you need to prove something, hash the file again and compare. If the two hashes match, the file is byte-for-byte identical to the one you hashed originally.
You can try step one yourself.
On Linux, run sha256sum yourfile.pdf.
On macOS, run shasum -a 256 yourfile.pdf.
In Windows PowerShell, run Get-FileHash yourfile.pdf -Algorithm SHA256.
All three give the same result for the same file, which is the point: anyone can check the hash with standard tools and doesn't need to trust a particular vendor's software.
Where hashing alone falls short?
This is the part that often gets skipped.
The hash can't recover the file.
If you hash a document and then lose it, the hash is useless as a copy. It can confirm a match only if you can produce the original. For anything you can't afford to lose, you still need your own backup.
It is byte-exact.
Opening a PDF and re-saving it, converting a Word file to another format, or even changing a file's metadata can change the hash while the visible content looks the same. A mismatch tells you the bytes differ. It doesn't tell you why or whether the difference matters. Careful teams hash the final, frozen version of a document and note which version it was.
A hash isn't encryption, and predictable documents can be guessed.
Hashing is deterministic. If a document is short or highly predictable, such as a standard form with a few fill-in fields, someone could hash candidate versions until one matches. For large or unique files this is not a practical concern, but it is a real one for tiny, guessable inputs. If that applies to you, consider adding secret random data to the file, for example inside a container you keep private, before hashing.
A hash alone says nothing about time.
A hash sitting in your own database proves nothing, because you could have computed it yesterday or edited the database today. A hash only becomes evidence of timing when it is recorded somewhere whose timestamp you can't rewrite.
It doesn't show who made the file.
A matching hash shows the same bytes existed. It doesn't show authorship, ownership, or originality.
Adding a trustworthy timestamp
The usual ways to fix the timing problem are a trusted timestamp authority, which issues timestamp tokens under the RFC 3161 protocol, or anchoring the hash to a public blockchain. In both cases only the hash is submitted, not the document. The record then shows that a particular fingerprint existed at or before a particular time, and the original file is never revealed.
How Certelo approaches it?
Certelo is built around this model. The SHA-256 hash is calculated locally on your device, and the original file isn't uploaded just to create the proof. The hash is then anchored to the Electra Protocol blockchain, which gives you a publicly verifiable record that the fingerprint existed at or before the recorded blockchain time.
Verification works in the other direction. You can check a record using the original file, the hash, or the certificate information. Because the hash is a standard SHA-256 value, you can also compute it yourself with the commands above and confirm it matches, so the comparison doesn't depend on trusting Certelo's own software.
The limits described earlier apply here as well. Certelo gives you evidence of a fingerprint at a point in time. It doesn't establish authorship, legal ownership, or that a work is original, and whether the record is admissible or persuasive depends on the jurisdiction and the dispute. The original file also stays with you, so keeping it safe is your responsibility.
When hashing instead of storing makes sense?
It suits situations where you need integrity and timing evidence but not a hosted copy. Examples include research records, invention notes, draft manuscripts, design files, compliance snapshots, and audit logs. It is a poor fit when you need the service to keep, search, or share the document for you. In that case you need storage, with or without a hash on top.
Frequently Asked Questions
Can I store a hash instead of a file?
You can, if all you need is to check later that a file hasn't changed. You can't use the hash to get the file back, so the original must be kept somewhere.
Can the original document be recovered from its hash?
No. SHA-256 is designed so that working backwards is not computationally feasible. The exception is guessing: if the document is short or predictable, an attacker can hash likely candidates and compare.
Is hashing the same as encryption?
No. Encryption is reversible with a key. Hashing is one-way and has no key.
What happens if the file changes slightly?
The hash changes completely. That's useful for detecting tampering, but it also means harmless changes like re-saving or reformatting will cause a mismatch.
Does a hash prove I created the document?
No. It shows that a file with those exact bytes existed. Authorship is a separate question that needs other evidence.
Does Certelo upload my original file?
No. The hash is generated locally in your browser, and the file isn't uploaded merely to create the proof.
Can I verify a record independently?
You can recompute the SHA-256 hash with standard tools and compare it with the recorded value. The blockchain record can also be checked through the Certelo verification process.
Summary
Hashing a document gives you a compact, private way to check integrity later, and combined with a trustworthy timestamp it becomes evidence of timing too. It doesn't replace the file, so keep your original, know what the hash can and can't show, and hash the final version of whatever you care about.