Automatic classification, and why GDPR compliance depends on it
You cannot protect, report on or delete personal data you cannot find. How NomadVault classifies documents on your device before encryption, and what that gives you under the GDPR.
Most GDPR obligations share an unstated prerequisite: you have to know where the personal data is. The record of processing activities assumes you can list it. A subject access request assumes you can find one person’s data within a month. The 72-hour breach notification assumes you can say what was in the affected system. Storage limitation assumes you know what has outlived its purpose.
In a document platform, none of that is given. Personal data does not arrive labelled. It arrives as a CSV export someone dropped in a project folder, a support log with customer emails, a contract with a date of birth on page four. Encryption, which is the right answer to confidentiality, makes the problem harder: a server that cannot read documents cannot scan them for you either.
This article explains how NomadVault handles that, and why we think classification at upload time, on your device, is the only version of the feature that fits an end-to-end encrypted product.
What classification does
When classification is enabled for your tenant, every text-based file is scanned in the browser at upload, after any content cleaning and before encryption. The scanner looks for indicators of four categories, with patterns for both English and German documents:
- Personal data. Email addresses, phone numbers, national identifiers, dates of birth, and the field labels that usually sit next to them: name, address, date of birth, passport or ID number, insurance number and so on.
- Financial data. Card numbers, IBANs, bank details, and terms such as salary, income or account balance.
- Health data. Diagnosis, prescription, medical report, lab result, common disease names. This category is off by default because it is a special category under Article 9, and switching it on is a decision your data protection officer should make knowingly.
- Credentials. Passwords, API keys and tokens that have no business being in a shared folder at all.
Your administrator can add custom patterns, for example your own customer number format or a project code word. Each detected category is stored as a label on the file. The content it was found in stays encrypted and leaves the browser only as ciphertext, as always.
A user who opens the file’s properties sees the detected categories and can override them, adding a category the scanner missed or removing a false positive. Overrides are marked as manually set, and every change is written to the audit log.
Where the labels go to work
A label on its own is a hint. Its value comes from what reads it.
The data inventory report. Administrators can generate a list of every file with its classification and any scheduled retention or deletion date, and export it as CSV. This is the closest thing to an Article 30 record a document store can produce by itself: here is where personal data lives, in which folders, owned by whom, due for deletion when.
Workflows. Classification is a trigger. A tenant can define that a file classified as health data in a folder shared with external guests sends an alert to the compliance team, or that a file matching a custom pattern is moved to a restricted folder, or that credentials found in an upload raise a ticket. The rules run on labels, not content, so they work without anyone, including the automation, reading the document.
Retention. Folders can carry an expiry with automatic deletion of contents. Combined with the inventory, this is how storage limitation becomes a schedule instead of an intention.
Access reviews. The admin reports that show who has access to what can be read against the classification labels, so a review of “who can see health data” is a filter, not a project.
Why on the device, before encryption
The obvious place to run a scanner is the server, after upload, where it can see everything. For NomadVault that is not an option, and we think it should not be one for any product that claims end-to-end encryption.
Running classification in the browser means the plaintext never leaves the device for the purpose of being analysed. The server receives the label and the ciphertext, nothing in between. It also means the trade-off is explicit: the label is one of the few things the server can see about a document, alongside its size, timestamps and access list, and we say so in the security overview. A reviewer who asks “what does the provider learn from classification” gets a one-line answer: the category names, nothing else.
There is a second reason. A scanner that runs before encryption runs before the data is committed anywhere. A credential or a card number caught at upload can be rejected, flagged or moved before it ever sits in a folder forty people can open.
What it does not do
Classification is pattern matching, and pattern matching has edges. We would rather you know them than discover them in an audit.
- Text formats only. The scanner reads plain text, CSV, JSON, Markdown, logs, configuration and source files, up to the first half megabyte. It does not open PDFs or Office documents, and it does not read images. A scanned contract is invisible to it. Where sensitive data lives mostly in Office files, the custom patterns and manual overrides carry more of the weight, and the gap is on our list.
- False negatives exist. A pattern for IBANs will not recognise a bank account written in prose. A false negative means the file is unlabelled and no workflow fires. Classification narrows the search; it does not replace a human owner knowing what a folder is for.
- False positives exist too. A phone number in a vendor’s signature makes an invoice “personal data”. That is usually correct under the GDPR, which is the point, but it means the labels describe presence, not sensitivity.
- Labels are visible to the server. By design, so that reports and workflows can run. If the fact that a folder contains health data is itself something your tenant must hide from the operator, leave that category off and use custom patterns with neutral names.
- Files uploaded before classification was enabled are not labelled until they are re-uploaded or labelled by hand.
How this maps to the GDPR
For the compliance reader, the connections in one place:
- Article 30, records of processing. The data inventory gives you the location, category and retention of personal data inside the platform, exportable on request.
- Article 32, security of processing. Classification-triggered alerts and automatic moves are technical measures that act on sensitivity, and they are logged.
- Article 5(1)(e), storage limitation. Folder expiry plus the inventory’s retention column turn a principle into dated actions.
- Article 15, access by the data subject. Combined with search, labels narrow where to look for one person’s data within the deadline.
- Articles 33 and 34, breach notification. If an account or a share is compromised, the inventory tells you within minutes whether personal or special-category data was in scope, which decides whether and whom you must notify.
- Article 9, special categories. Health detection is opt-in, so processing it is a documented decision, not a default.
None of this makes a tenant compliant on its own. It gives the people responsible for compliance the one thing an encrypted store otherwise withholds from them: a map.
Questions?
If you want to discuss how classification fits your records of processing, or your data protection officer wants the detection rules and the report format, get in touch at hello@nomadvault.de.