SharePoint data classification: what it is and how to do it
Classifying data in SharePoint means recording what each file contains, such as payment card numbers, bank details, identity numbers or contract terms, so that protection, retention and access decisions can be made on content instead of on location. In SharePoint Online you do this with Microsoft Purview. Sensitive information types find the content, sensitivity labels record the result, and auto-labeling policies and on-demand classification apply it to files that are already stored.
Each of those has limits, and the limits decide what a first pass will miss. This page covers the tools, the limits Microsoft documents for them, a step-by-step path for a tenant, and what changes when content has been archived out of SharePoint, where the usual tools stop seeing the document.
Classification works at three levels in SharePoint
| Level | What it classifies | How it is done |
|---|---|---|
| Site or container | The site, team or group as a whole | A default sensitivity label for the container |
| Library | The files in one document library | A default sensitivity label for the library, which Microsoft documents a way to apply to files already in it |
| File content | What is inside each file | Classifiers that read the content, then labels applied by auto-labeling or an on-demand scan |
The first two record where a file lives. Only the third looks at what the file holds, which is why most "we need to classify our data" projects turn out to be about file content, and why the rest of this page is about it.
How Microsoft finds sensitive content in files
Microsoft Purview groups its detectors under Classifiers, which you reach from Information Protection, Data Loss Prevention or Data Security Posture Management in the Purview portal. The types, as Microsoft describes them:
- Sensitive information types match patterns of sensitive information, such as social security, credit card or bank account numbers. You can use the predefined types or create your own.
- Named entities use dictionaries and patterns to detect person names, physical addresses, and medical terms and conditions.
- Exact data match identifies sensitive content by matching exact values.
- Trainable classifiers recognise a kind of document. Microsoft provides pretrained ones, and you can train custom ones by supplying samples.
- Credentials detect credentials from supported services and environments.
- Document fingerprinting recognises content that is a variation of a template.
- Optical character recognition lets Purview scan images for sensitive information, which extends classification to images without separate classifiers.
Microsoft describes the newer Classifiers experience as being in preview alongside the existing experiences, so the portal navigation may differ from older guides.
How labels get onto files that are already in SharePoint
Writing a detector does not label anything. A label reaches a file already in SharePoint in one of two ways: an auto-labeling policy, or an on-demand classification scan.
Auto-labeling policies
Service-side auto-labeling runs in the Microsoft 365 service, not in the user's Office app, and covers SharePoint, OneDrive and Exchange. What Microsoft documents for SharePoint and OneDrive files:
- PDF documents and Office files in Word (.docx), PowerPoint (.pptx) and Excel (.xlsx) format are supported.
- Files can be labeled at rest, before or after the policy is created.
- A file that is part of an open session, because someone has it open, cannot be labeled.
- Attachments to list items are not supported and are not labeled.
- Microsoft lists service-side auto-labeling under Microsoft 365 E5, Microsoft 365 E5 Compliance, Microsoft 365 E5 Information Protection and Governance, and Azure Information Protection Premium P2.
Run the policy in simulation mode first. Simulation supports up to 20,000,000 matched files, and if a policy matches more than that you cannot turn it on until you narrow it and simulate again. Microsoft also notes that simulation results can differ from what happens when the policy is live, because simulation shows the result of a single policy.
On-demand classification
Microsoft Purview on-demand classification identifies and classifies sensitive content in historical data in SharePoint, OneDrive and endpoints. Microsoft describes it as extending classification to files that have not been classified or modified for a long time, or that need updated classification.
How it behaves, from Microsoft's documentation:
- You define a name, the scope (all SharePoint sites and OneDrive accounts, or chosen ones), the classifiers, a last modified date range, and the file extensions.
- By default the scan includes items created or modified within the past year, across all supported extensions and every classifier configured in the tenant. You can choose up to 50 specific classifiers at a time.
- An estimation runs first, showing progress, items found and estimated cost, and you start classification after reviewing it.
- Each scan can process up to 150,000 locations and 100 million files.
- You need the Compliance Administrator role group to run a scan, and a Content Explorer viewer role to see results.
- It uses pay-as-you-go billing or per-user licensing for Purview capabilities, or both.
- Content Explorer, which the current Learn overview calls Data Explorer, updates within seven days of scanning.
- If you select specific classifiers, only those classifiers have their results updated for the scanned files.
Where classification stops: what a first pass misses
These come from Microsoft's own documentation, not from a competitor's marketing. Each is a reason a SharePoint estate can look classified and not be.
Older files that have not changed. Microsoft's own checklist for files that were not labeled says: "Auto-labeling only evaluates content that was created or modified after the SIT was created or last modified. To classify older unchanged files, run on-demand classification." Those are the files most likely to be forgotten, and the on-demand scan's default window is the past year, so widen the date range deliberately.
Formats outside the supported list. Auto-labeling for SharePoint lists PDF and the modern Word, PowerPoint and Excel formats as supported. Microsoft's failure table includes codes for file types and extensions that cannot carry a sensitivity label, and for those no change to the policy will help.
Protected files. A file protected by external encryption, such as a password, cannot be labeled until the protection is removed.
Files in use. A file that is open in a session cannot be labeled while it is open. Microsoft's failure table shows locked and checked-out files being retried automatically.
Sites over their storage quota. Auto-labeling reports a QuotaExceeded failure when the site's storage quota is exceeded, and Microsoft's action is to free up storage or increase the quota and rerun the policy.
A label that already outranks the policy. If a file already carries a label of equal or higher priority, the policy does not override it. Microsoft treats that as by design.
Every one of these surfaces as a code against the file in the Purview portal. The full list, with what each one means and which ones fix themselves, is in Purview auto-labeling failure codes for SharePoint files.
A practical path to classify a SharePoint tenant
- Decide the categories before the tool. List what you need to find, such as payment cards, bank accounts, identity numbers, health information or contract terms, and give each a severity. A tool configured without this produces a count nobody can act on.
- Choose the detectors. Use the predefined sensitive information types where they fit, and write your own where your identifiers do not.
- Simulate before you turn anything on. Read the simulation results, remembering they show one policy at a time.
- Turn on the policy for new and changed content, then run an on-demand scan for what is already there, with the date range widened beyond the default year.
- Review the results in Content Explorer. Allow for the delay of up to seven days after a scan.
- Work the failure list. Most codes retry on their own. A few need an action from you.
- Run the on-demand scan again when a detector changes. Auto-labeling will not revisit older unchanged files on its own.
- Decide where inactive content lives before you classify it. Content you archive moves out of the place these tools look.
When the content has been archived
When Squirrel archives a file, the document moves, encrypted and compressed, into your own Azure storage, and SharePoint keeps only a small pointer page in its place. Anything that looks at SharePoint sees the pointer page, not the document. The obligation to know what that content holds has not moved with it, so an archived estate needs a tool that can reach inside the archive.
That is the gap Forage fills.
How Forage classifies SharePoint content, live and archived
Forage is the records half of Squirrel. It reads the documents in SharePoint Online and in the Squirrel archive, including the words inside PDFs, Office files, emails and scanned pages, and makes them searchable in one place. It builds its inventory from the records Squirrel already keeps, so SharePoint is not crawled a second time.
Where content classification is enabled, which is an add-on to the search foundation, Forage recognises what documents contain: bank details, payment card numbers, tax and company identifiers, passport numbers, health information and contract terms. It does this with transparent rules built from phrases, patterns, checksum validation and proximity. Every finding keeps the words that produced it, and the same document always gives the same result.
What an administrator can do with it:
- Write, version and test the rules in the portal. Before a rule goes live, a test run shows the most it could match, a sample checked with the real rule, and a projected range across the documents read so far.
- See the result label by label: how many documents contain each kind of content, how serious it is, and in which sites, and open the documents behind any label to read the matched text in context.
- Search by label, with or without words, together with the other search filters, across live and archived content in one search.
- Account for every file. Each file is read, excluded by type or by policy, or recorded as failed with the reason, and coverage is reported against the sites you chose.
- Report on it. The Sensitive content and Labels by file reports export as PDF or a spreadsheet, or are emailed on a schedule you set.
- Record who did what, in an audit log that no one using Forage, administrators included, can change or remove an entry from.
What Forage does not do
Forage's limits are stated as plainly as its features.
- It requires Squirrel. Forage is offered only to organisations that use Squirrel, and it covers SharePoint Online and the content Squirrel has archived from it. OneDrive, Exchange, Teams and other systems are not in scope.
- It does not apply Microsoft labels. Forage's labels are its own content labels, recording what a document contains. It never applies, changes or removes a Microsoft label, and it is not a replacement for Microsoft Purview.
- Its rules match exactly. They match the words and patterns written into them, with checksum checks and nearby-word conditions where a rule asks for them, and make no judgement about meaning.
- It reads the current version only. Earlier versions of a document are not read.
- Findings cover what has been read. Reading a large estate takes time, and the Coverage page shows how much has been read so far.
- Protected files are recorded, not read. A password-protected or encrypted file is recorded as protected and its content is not searchable.
- Retention and disposal are not yet available. Forage deletes nothing today.
The wiki covers classification, writing content rules, searching by label and the full limitations. To see it read your own estate, email sales@smikar.com.
Frequently asked questions
What is data classification in SharePoint?
It is recording what a file contains so that decisions can be made on content rather than location. In SharePoint Online that means Microsoft Purview detectors finding the content, and sensitivity labels recording the result on the file.
Does SharePoint classify data automatically?
Not by default. Labels reach files already in SharePoint through an auto-labeling policy or an on-demand classification scan, and both need to be configured. Microsoft lists service-side auto-labeling under E5-tier licences.
How do I find sensitive data that is already in SharePoint?
Run an on-demand classification scan with the classifiers you need, then review the results in Content Explorer. Widen the last modified date range, because the default covers only items created or modified within the past year.
Why were my older SharePoint files not labeled?
Auto-labeling evaluates only content created or modified after the sensitive information type was created or last modified. Older unchanged files need an on-demand scan.
Can I classify files that have been archived out of SharePoint?
Tools that look at SharePoint see the pointer page left behind, not the archived document. For content archived by Squirrel, Forage reads the archive copy and classifies it where content classification is enabled.
Does Forage apply sensitivity labels?
No. Forage records what documents contain using its own content labels. It never applies, changes or removes a Microsoft label, and it is not a Purview replacement.
Can I use Forage without Squirrel?
No. Forage requires Squirrel and is offered only to organisations that use it.
Related reading
- Purview auto-labeling failure codes for SharePoint files - every code, what it means and which ones fix themselves.
- Sensitivity labels vs retention labels in SharePoint - the two label types admins confuse.
- Retention policies vs retention labels in SharePoint - policies apply to locations, labels to items.
- Microsoft Purview explained - the pieces of Purview and where each fits.
- Purview retention archiving for SharePoint - the Message Center change that archives files under retention policies.
- Data loss prevention in SharePoint Online - using classification to stop sensitive content leaving.
- Prepare SharePoint for Copilot - why unlabeled content is a Copilot risk.
- Forage: records management for SharePoint and the Squirrel archive
Mark Smith co-founded SmiKar Software in 2015 and has spent the past decade helping organisations solve Microsoft 365 data management challenges. He works with the SmiKar team to build solutions for SharePoint archiving, storage optimisation, governance and compliance, supporting customers from growing businesses through to Fortune 500 enterprises.
More about SmiKar


