The Data Classification Problem Nobody Wants to Solve

Data classification in 2026 stands as the security work that everyone agrees needs to happen and that nobody wants to do. The work takes months, the work costs money, the work produces a taxonomy nobody uses, and the work does…

Dark cinematic editorial image for The Data Classification Problem Nobody Wants to Solve - abstract cyan and electric blue digital composition in deep black, hacker aesthetic, no text no logos

5 MIN READ

Most companies have 10 plus petabytes of data scattered across 200 plus data stores, and the classification of the data sits somewhere between “we know some of it is sensitive” and “we have not classified the rest.” The 2026 state of the data classification problem is a problem that almost nobody has solved, and the gap is where the next regulator letter, the next AI hallucination, and the next DLP alert are all going to come from.

Snowflake, BigQuery, Databricks, Redshift on the warehouse side. Salesforce, HubSpot, Zendesk, Workday, the 200 plus other SaaS applications the typical company runs. On prem databases, file shares, the email, the document stores, the collaboration tools, the chat tools. The data inside those stores ranges from public marketing materials to customer PII to financial records to trade secrets. The classification in the typical company sits at “we know some of it is sensitive, we have not classified the rest,” and the AI models the company is trying to ship in 2026 are being trained on the part nobody has classified.

Why the work matters more than ever

Compliance, the first pillar. GDPR Article 30 records of processing, CCPA data inventory requirements, HIPAA risk analysis, PCI cardholder data inventory, the SOX financial data handling, the sector specific overlays in healthcare (HITRUST), financial services (FFIEC, NYDFS 500.11), and the EU (DORA). The CISO who does not know what data the company holds cannot answer the regulator when the regulator writes the letter, and the regulator is writing more of these letters each year. Security, the second pillar. The DLP, the access control, the encryption, the data masking all depend on knowing what data those controls are protecting. Security controls applied to all the data run expensive and slow. Security controls applied to the sensitive data run focused and effective, and the difference is the classification. AI, the third pillar. The customer service AI, the document search AI, the code completion AI, the RAG pipelines the company is shipping all need to know which data the model can use, which data the model cannot use, which data the model needs explicit consent for. Microsoft Copilot, Google Workspace Gemini, and the Salesforce Einstein rollouts are all blocking on this question, and the CISO who has not classified the data cannot ship the AI deployment.

Why the typical company has not done the work

Scope problem, the first. The data classification project starts as a 12 month project, the project gets stuck on the long tail of the data stores, the project never finishes, and the project gets renamed (data governance initiative, data trust framework) and quietly shelved. Taxonomy problem, the second. The CISO builds a perfect 5 level taxonomy (public, internal, confidential, restricted, regulated), the 5 levels amount to too many for the data stewards to agree on, the taxonomy gets debated for 6 months, and the taxonomy never gets used. The Microsoft, Google, and AWS reference taxonomies are 4 levels at most, and the ones that ship tend to be 3. Automation gap, the third. The CISO tries to classify the data manually, the manual classification takes forever, the manual classification runs wrong half the time, and the manual classification gets abandoned. The cloud providers have been closing the automation gap since 2023, and the CISO who has not looked at Microsoft Purview, AWS Macie, and Google Cloud Sensitive Data Protection is the CISO still trying to do this with a spreadsheet and a quarterly review.

How to actually do the work

Start with the data inventory rather than the classification. The inventory answers “what data do we have, where does it live, how much of it is there, who can access it,” and the inventory becomes the foundation everything else sits on. AWS Macie, Google Cloud Sensitive Data Protection, Microsoft Purview, and the open source options (Apache Atlas, DataHub, OpenMetadata) all produce the inventory with a few weeks of configuration. The classification without the inventory amounts to a classification nobody trusts. Pick a 2 or 3 level taxonomy rather than a 5 or 6 level one. Public, internal, confidential, restricted amounts to the 4 level taxonomy Microsoft Purview ships with, and most companies can collapse it to 3. The 2 or 3 level taxonomy amounts to the taxonomy the data stewards can actually apply, and the data stewards applying it consistently amounts to the only thing that ships the project. Automate the classification last, with the cloud provider tools (Microsoft Purview covers 100 plus data source types, AWS Macie handles the S3 inventory, Google Cloud DLP scans the BigQuery and Cloud Storage estates) and the pattern matching for the structured data, the content inspection for the unstructured. The company that does the inventory, picks the simple taxonomy, and automates the classification amounts to the company that finally finishes the work.

Abstract data classification as glowing cyan nested categories on a dark navy surface, dramatic chiaroscuro lighting from above
Data classification in 2026: 3 reasons it matters (compliance, security, AI), 3 reasons the typical company has not done it (scope, taxonomy, automation), 3 moves to do it.

The bottom line

Inventory first, simple taxonomy, automated classification. The company that does the work ships the AI, runs the DLP, and answers the regulator in the same week, and the company that has not done the work stays stuck on all three.

Sources & Further Reading

All claims in this article are sourced from primary documentation, vendor advisories, and reputable security researchers.

Spotted an error? Email the editor. Corrections are issued with a visible correction note.

Editorial standards. Every article on humanrequired.org is reviewed by a human editor before publication. AI may assist with drafting or research; final editorial control is human. Read the full standards.

Continue reading