EmberHound Personal Data Discovery
Find personal data across every enrolled company device.
EmberHound searches files, documents, and images for personal information, then shows your team where it is stored, what presents the greatest exposure, and where to look when a subject access request arrives. Raw files stay on the device.
- Windows, macOS, and Linux
- On-device scanning
- Raw files are never uploaded
- Masked findings only
The problem
Personal data spreads through everyday work.
An HR spreadsheet pulled for one review meeting. A customer list exported to help with a mailshot. A scanned sick note saved to a desktop. A support ticket copied into a document so somebody could work on it offline. None of it is anybody being careless - it is just how work gets done.
Interviews, policies, and a manually maintained inventory tell you what your organisation intends to hold. They do not tell you what is on the laptops.
Exports and downloads
Reports pulled for one task, saved to Downloads, and never removed.
Documents and attachments
HR records, scanned forms, screenshots, and email attachments saved locally to work on.
Archives and old folders
Zipped backups and operational folders from projects that ended years ago.
Two jobs
Related, but not the same task.
They use the same findings. They answer different questions, and conflating them is how "GDPR compliance" becomes a word that means nothing.
Personal-data discovery
Ongoing. Understand which categories of personal information exist, where they are stored, and which devices or locations need attention. This is the work that builds the picture.
DSAR search
Reactive, and on a clock. Locate the findings associated with a particular person's identifiers so your team knows where to look. This is the work that answers one request.
What you get
Four things, from the same scan.
See where personal data is stored
Identify supported categories of personal information across enrolled endpoints, by category, device, and location.
Find forgotten or unnecessary copies
Surface data in downloads, exports, archives, and images that will not appear in a manually maintained inventory.
Locate information for a person
Cross-reference known identifiers against existing findings when a subject access request arrives.
Keep evidence of the work
Scan history is retained, and findings and responses export as records of what you searched and what you found.
Detection
What EmberHound looks for.
The core pack covers contact and identity details and government identifiers, and separately the categories that carry extra conditions for processing under the regulation.
Contact and identity
Email addresses
standard address format
Telephone numbers
international dialling formats
Postal addresses
street-suffix patterns, including Strasse, Rue, and Chemin
Dates of birth
day/month/year and ISO forms
Names
matched where a title such as Mr, Ms, Dr, or Prof precedes them
National and government identifiers
UK National Insurance numbers
format-validated
UK passport numbers
nine-digit format
UK driving licence numbers
DVLA format
UK Unique Taxpayer Reference
ten-digit format
IBAN
validated with an IBAN checksum
Article 9
Special category data
Detected by dedicated rules, reported under their own subtype, and switchable per category for your organisation.
Article 10
Criminal offence data
Handled separately from Article 9, because it is a different regime with its own conditions for processing rather than a further special category.
We do not publish which terms or fields those two sets of rules match on. Setting that out on a public page would tell anybody who wanted to avoid detection how to, and it is not the right place to discuss scanning for information about somebody's health, faith, or criminal record. We will walk through it with you directly.
Two honest limits. Names are matched where a title precedes them, so a bare first and last name in a paragraph will not be picked up on its own. IP addresses and other online identifiers are not detected at all.
Regional detection
National identifiers do not share a format.
A rule that finds a UK National Insurance number will not find a German tax ID. Country packs add the formats and the checksums for a particular country, and every finding records which country rule produced it.
Two country packs are published today. Others are not covered yet, though you can extend detection with custom rule overlays. A country pack tells you what a string is - it does not determine which law applies to your organisation.
United Kingdom
In the core pack
National Insurance number, passport, driving licence, and Unique Taxpayer Reference.
Germany
Country pack
Personalausweis, social security number, and Steuer-ID with its checksum.
France
Country pack
Carte nationale d'identite, NIR with its checksum, and SIREN.
How it works
From agent install to exported evidence.
- 01
Deploy the endpoint agent
Push it through your existing MDM, or hand somebody an enrolment code.
- 02
Choose what to search for
Select the personal-data policy, add country packs where relevant, and set which locations are in scope.
- 03
Scan locally
Matching runs on the device. The file itself never leaves it.
- 04
Review findings
Category, device, file, masked context, confidence, and risk - grouped and deduplicated.
- 05
Prioritise remediation
Assign an owner, set a status, work down by risk. Findings carry an SLA due date.
- 06
Rescan and export
Check whether exposure actually changed, then export the record.
EmberHound never deletes, moves, edits, or redacts the files it scans. Remediation is your team's action, tracked in the platform.
Findings
See enough to investigate without collecting the file.
A finding has to tell somebody where to go and what they will find. It must not become a second copy of the personal data you are trying to get under control. So the preview is masked, the value is stored as a hash, and the rest is metadata.
Findings locate the match inside the file itself: the page of a PDF, the sheet and cell of a spreadsheet, the path within a JSON document.
- Data category and subtype
- Masked preview, never the value itself
- Device, file path, and file type
- Where inside the file: page, sheet, row, or column
- Confidence score, and what adjusted it
- Risk level and exposure level
- Which country rule applied, where one did
- How many times it has been seen, across how many scans
- Status, assigned owner, and resolution note
- Whether OCR was used to read it
Signal
Focus on credible findings, not every possible match.
A nine-digit number is a passport, a SIREN, an invoice reference, or nothing at all. Six controls separate them, and each finding records which applied to it.
Validation rules
IBAN, the German tax ID, and the French NIR are checksum-validated. Each finding records whether the checksum passed.
Required context
Weak patterns are not reported alone. Special-category rules mostly look for a labelled field and its recorded value, not a word in passing.
Negative matching
The French national ID rule actively excludes invoice, order, ticket, and asset references that share its digit pattern.
Confidence scoring
Each rule carries its own threshold, and every finding records the raw score plus the adjustments applied to it.
Deduplication
A fingerprint groups the same value across scans and devices, so one record in fifty copies of an export reads as one thing.
Suppression rules
Dismiss a finding or exclude a path from future scans. A suppression needs a written reason of at least 10 characters, and it is audited.
None of this eliminates false positives. It moves the credible findings to the top of the list.
Find personal data inside images and scanned documents.
A scanned passport, a photographed form, a screenshot of an HR record. Ordinary file search reads none of them, because there is no text in the file to read.
With the OCR add-on enabled, text is extracted from supported images and image-based documents before the same rules run, and each finding records whether OCR produced it.
OCR, mailbox scanning, and drive scanning are add-ons rather than part of a base plan. Ask us about add-on coverage.
- Applies the same rules to extracted text as to ordinary files
- Every finding records whether OCR produced it
- OCR confidence is recorded so low-confidence reads can be filtered
- Extraction happens on the device, like the rest of scanning
Mailbox scanning is not a cloud integration.
The mailbox add-on extends what the endpoint agent reads. EmberHound does not connect to a mail platform and pull messages from it. If that is what you need, talk to us about your estate before buying.
Subject access requests
When a subject access request arrives, know where to look.
The hard part of a DSAR is rarely the paperwork. It is establishing where a particular person's information actually sits across an estate nobody has a complete picture of. If your endpoints have been scanned, that search is a lookup rather than an expedition.
Identifiers you can search on
- Email address
- Telephone number
- Name
- Postal address
- Date of birth
- National identifier
- IBAN
- 1
Log the request and verify identity
Record the request and how identity was confirmed - email confirmation, document upload, knowledge-based checks, a trusted introducer, in person, or account ownership.
- 2
Track the deadline
The due date is held on the request, along with any extension and the date the subject was told about it.
- 3
Search by identifier
Identifiers are normalised and hashed with an organisation-specific pepper, then matched against finding hashes. Raw identifiers are not stored in the clear.
- 4
Review what matched
Matches are clustered and scored into confidence bands. A reviewer includes or excludes each cluster with a reason code, and the decision is recorded against them.
- 5
Produce the response
An Article 15 response document sets out the personal data, the purposes and legal basis for each record group, an Article 15(1)(c) recipients table, and the subject's rights.
- 6
Deliver and close
Release through the subject portal, encrypted email, post, or API. The pack records who finalised it, who released it, and when it was delivered.
Where EmberHound stops
It covers the request from intake through to delivery, with audit history throughout. Two things it does not do, deliberately: it does not fetch the underlying documents, because they stay on the endpoint, and it does not redact them. It records the decisions your reviewers make about relevance, exemptions, and third-party information. It does not make those decisions for you.
A list of file locations is not a completed DSAR response. It is the part that used to take the longest.
On the deadline
Organisations normally need to respond without undue delay and within one month. That period can be extended where a request is complex or where somebody has made a number of requests, and the start point can be affected by needing to confirm identity or seek clarification.
It is not a flat 30 days, and it is not the 72 hours that applies to notifying a personal data breach. Those are different obligations with different clocks.
On the search
EmberHound helps your team search supported endpoints consistently and retain evidence of where it looked and what it found. Your organisation remains responsible for deciding the appropriate scope of each request.
A scan is evidence that a search was carried out and how. It is not, by itself, proof that the search was reasonable and proportionate for that particular request.
Records of processing
A record of processing that earns its keep.
The ROPA Workspace holds every Article 30 field, keeps entries under review and sign-off, and feeds signed-off activities into the Article 15 response document. A scan can't tell you your legal basis, your purposes, or your retention period - those are decisions, not facts on a disk. What it can do is show you personal data in places no entry accounts for. Write the record; use the findings to check it against reality.
GDPR
Support GDPR work with evidence from actual data discovery.
Scan results are evidence about storage. That evidence feeds personal-data inventories, data-mapping exercises, retention reviews, exposure remediation, security reviews, internal audits, and the accountability record behind all of them.
How this maps to the Articles
A scanner does not fulfil an Article. It produces evidence that helps you meet an obligation you hold.
Article 5 - Principles
Supports minimisation and accuracy work by showing what is actually stored, and supports accountability by keeping the record of what you found and did.
Article 15 - Right of access
Helps locate information relevant to a request, and produces the response document and delivery record.
Article 16 - Rectification
Supports rectification searches and records the work. EmberHound does not correct the underlying file.
Article 17 - Erasure
Supports erasure-location searches, tracks the work, and can issue a completion certificate. Deleting the file is your action.
Article 30 - Records of processing
Provides the register itself, with the Article 30 fields and a review cycle, plus scan evidence to check the entries against. Completing it does not by itself discharge the obligation.
Article 32 - Security of processing
On-device scanning, masked findings, tenant isolation, and audit history are evidence of technical measures.
EmberHound supports personal-data discovery, DSAR searches, and GDPR accountability work. It does not provide legal advice, guarantee that every relevant record will be found, or make an organisation compliant by itself.
ICO guidance
Security
Search personal data without collecting the underlying files.
A tool that found all your personal data by copying it somewhere central would have made the problem worse. Scanning runs on the endpoint and the files stay there.
One exception worth naming: a disclosure pack you choose to produce and deliver does contain the response document your team prepared. That is the point of it, and it is created deliberately rather than as a side effect of scanning.
Read the Trust Centre for the architecture and data-handling detail.
- Matching runs on the endpoint. The platform never reads your file system.
- Raw files stay on the device. Only a masked preview, a hash, and metadata leave it.
- DSAR identifiers are hashed with a pepper held for your organisation alone.
- Encrypted in transit and at rest.
- Organisation-level isolation, enforced in the database as well as the application.
- Role-based access, and every action written to an audit log.
Use cases
What teams run it for.
Map personal data across endpoints
Understand which categories of personal information are actually present on enrolled devices.
Investigate a subject access request
Move from the identifiers you were given to the devices and files your team needs to review.
Find unnecessary local copies
Surface exports, downloads, and archives that no longer have a business purpose behind them.
Review sensitive-category exposure
Article 9 and Article 10 data are covered and gated separately, so each can be looked at on its own terms.
Validate data maps and inventories
Compare what your records say you process against what the devices show you are storing.
Watch whether exposure is improving
Use scan history to see whether data your team dealt with has come back.
Frequently asked questions
From our blog
Practical GDPR guidance for IT administrators and security teams.
Related resources
- What is a DSAR and how do you respond within GDPR's deadline?
A practical guide to handling data subject access requests on time.
- GDPR Article 30: Build a ROPA Without a Spreadsheet
How to maintain processing records using automated data discovery.
- Data Breach Notification Under GDPR: The First 72 Hours
Step-by-step obligations from discovery to regulator notification.
- How to Prepare for a GDPR Audit: 10-Point Checklist
A practical audit preparation guide for IT administrators.
Find out where personal data is being stored.
Run your first scan on one enrolled device and see the categories, locations, and exposure EmberHound identifies, without uploading the underlying files.
No card details required. One device, one user.
