Back to resources

Data Ready for AI Anywhere – Knowledge Still Protected

Every run was permitted, but nobody approved the map created by combining the data
Figure 1. Strava’s global heatmap. Every trace on it was permitted on its own, and together they drew a perimeter nobody meant to publish.

In January 2018, a twenty-year-old Australian student named Nathan Ruser was browsing a heatmap Strava had just published, built from roughly 13 trillion GPS points logged by people out running and cycling. Out in the Syrian desert, he noticed running routes glowing in the dark.

They belonged to soldiers. In places where nearly the only Strava users were visiting personnel, the map quietly traced the perimeters of forward operating bases, the supply roads between them and the patrol routes in and out. Militaries spent the following week rewriting policy.

Nobody had leaked anything, because every upload was a jog, permitted and unremarkable on its own. Nobody approved a map of base perimeters either, and nobody could have, because it did not exist until the individual runs were joined together.

The risk is not access. It is inference.

Enterprises are about to run that same play on purpose, at machine speed, and mostly with the best of intentions.

Take the fraud model or churn agent almost every business wants. It reads transactions, devices, channels, support calls and payment behavior, then tells you which customer is at risk while you can still act. To find the one, it has to profile all of them, and every field it touched it was allowed to touch. Nobody approved the resulting profile, though, and nobody could have, because it did not exist until the agent joined the dots.

Traditional data security asks who can access this table, who can query this field, and who can download this record. AI adds a question that none of those can answer. What can a model work out by combining data it was already permitted to see?

The quieter risk is standing still

There is a second risk, and it is the more expensive one, though it never announces itself as a security incident. It turns up quietly on the revenue line instead.

When teams cannot see inference risk clearly, they do the safe-looking thing and slow down, adding longer reviews, tighter limits and a few more nos. It feels careful, which is exactly what makes it so easy to miss that it is not.

Illustration showing the business cost of slowing AI programs despite investing in advanced capabilities
Figure 2. Stuck in second gear. The cost of waiting never files an incident report, but EMA’s numbers on delayed AI projects show it clearly enough.

Put it in money and the arithmetic gets uncomfortable. A fraud model sitting in review for a quarter is a quarter of fraud losses you quietly absorbed while it waited, and if a competitor ships their service AI two quarters ahead of you, they are the ones who set the customer’s expectation.

Teams only slow down because they cannot control the exposure, so give them a way to use the data safely and the reason to wait simply disappears.

Cloudera governs the platform. Protegrity protects the data in use.

Cloudera customers already have a strong foundation to build on. Cloudera Shared Data Experience (SDX) provides unified, cross-cloud governance, metadata management, and lineage, while Apache Ranger controls who can access data and what they can do with it. Both do their job well, which is worth saying before we talk about where the job ends.

None of it is wrong. It all just stops in the same place, which is easiest to see if you picture a relay. Discovery tells you where sensitive data lives, but it cannot follow that data into a prompt. DLP watches the exits, but it cannot see what a model reasons about once it is inside. Platform controls do real work, so long as the work stays inside the platform. Then the data lands in a model, an agent’s memory or a knowledge graph, and whatever protection it was carrying is all it has left, which is usually nothing at all.

Nobody on this track has a bad leg, and every tool does exactly what it was built to do. The baton just gets dropped in the open space between the last control and the model.

Relay illustration showing discovery, DLP and platform controls each running their leg before data reaches the model
Figure 3. Three clean handoffs and a dropped baton. Discovery, DLP and platform controls each run their leg, then the data enters the model and nobody is left holding it.

That last leg is ours. Protegrity applies protection to the data itself, so sensitive values stay protected as they are queried, shared and reasoned over, with policy deciding who sees clear text and under what conditions. SDX governs the door, and we protect what is inside the room, even after someone has carried it out.

Cloudera named Protegrity its ISV Partner of the Year for 2026, and the integration is native and certified across Cloudera Data Platform, Cloudera Data Engineering and Cloudera Data Warehouse, with support for Iceberg, Spark, Hive, NiFi, Kafka, Impala and Ranger.

Diagram showing protection that travels with sensitive data across AI workflows
Figure 4. The three pillars of protection for AI, covering knowledge protection, leakage prevention and attestation.

Together we move customers from restricting sensitive data to making it safe to use.

“Our customers don’t want to lock more data away; they want to safely deploy their most valuable, previously restricted data into AI pipelines. Cloudera’s Shared Data Experience (SDX) establish the unified governance baseline across the entire data lifecycle, and by embedding Protegrity’s data-centric protection natively into our hybrid platform, organizations can accelerate AI time-to-market in a governed and secure way.”

— Carlos Eduardo Zorzin, Partner Solutions Engineering Manager, Cloudera

“The moment a value leaves the table and enters a prompt, it is out of reach of everything that was guarding it. So we stop protecting the place and protect the value. It arrives at the model already tokenized, it comes back only to the person policy allows, and every reveal is a decision somebody can point to later.”

— Muneeb Hasan, Senior Solutions Engineer, Protegrity

What it looks like in practice

Take secure text-to-analytics on Cloudera’s SQL AI Assistant. A user asks a question in plain language in Cloudera Data Explorer, and before it reaches the model, Protegrity’s semantic guardrails scan the prompt, tokenize any PII and block prompt injection. The query then runs within Cloudera’s lakehouse engines where the Protegrity protector intercepts it at the execution engine, Hive or Impala, and decides per user, in memory, whether that person sees tokens or clear text. On the way back out, the response is scanned once more.

Secure text-to-analytics process on Cloudera showing protection from prompt through query and response
Figure 5. Secure text-to-analytics on Cloudera, from a plain-language question in Cloudera Data Explorer to an answer that comes back with the sensitive values still tokenized.

The analyst gets a useful answer, the sensitive values never reach the model, and every reveal is deliberate and logged. That is governance happening during the interaction rather than a report about it afterward.

Two things make this practical. The first is that tokenization is format-preserving and vaultless, so schemas, joins and BI tools keep working while models run on protected data without code changes. The second is that there is only ever one dataset, tokenized at rest and unprotected in memory at the final mile, which means no duplicate masked tables and no drift. The same dashboard can happily show a branch manager real values and an offshore agent nothing but tokens.

Because enforcement runs natively in the execution engine, there is no gateway hop on query workloads. That is how a top-5 global bank runs on protected data at 27 billion transactions a day across more than 100 countries, reporting 126% ROI within eight months.

Anywhere includes across borders

If protection lives inside the data, then the data stays safe wherever it goes, whether that is any cloud, on premises, the edge, or the one that matters most for regulated businesses, across borders. AI comes to the data instead of the data being extracted into a cloud-only silo, so the residency posture stays intact.

Picture a bank in five markets that wants one shared fraud model. Fraudsters do not stop at borders, but residency rules stop the data, which is where most of these projects quietly die. Protect the data itself and the protected version travels while the real values stay home.

Protected data crossing residency boundaries while clear sensitive values remain in their home market
Figure 6. Protected data crossing residency boundaries, so the model learns the pattern while the real values stay home.

That gives you one strong fraud model working across five markets instead of five weak ones fighting alone, and the extra fraud you catch is the business case.

A Global Systemically Important Bank (G-SIB) deployed this integrated Cloudera and Protegrity architecture, across on-premise and AWS cloud protecting 25+ core sensitive fields, unlocking 70% of cloud use cases that would otherwise have stayed blocked. IDC expects 80% of Asia’s largest enterprises to prioritize this kind of sovereignty by 2028.

Start here

Take your highest-value, most-restricted use case, the one stuck in review, and protect the data itself rather than the space around it. Deployment uses Cloudera Manager parcels onto your existing cluster nodes, with no new infrastructure, and for an existing Cloudera customer production is realistically four to eight weeks.

Watch the AI Workbench demo to see protection applied from prompt to query to response.

Explore the Protegrity and Cloudera integration to review the architecture, supported services, deployment options and policy enforcement at Protegrity and Cloudera.

Meet us at EVOLVE NYC on 27 October, where Protegrity joins Cloudera, IBM and Phoenix AI on a panel on Data and AI Anywhere. Book an architectural mapping session with joint Cloudera and Protegrity engineers while you are there.