Protegrity & Google Cloud Platform
Live Demo
View DemoProtegrity Native
Protegrity Cloud Protect uses cloud-native components deployed within Google Cloud to bring field-level protection closer to the applications, pipelines, and managed data services processing sensitive information.
Integration type
- Data Warehouse
Partner
Yes
Supported platforms
- GCP
overview
Google Cloud gives data teams managed services for ingestion, processing, analytics, and AI, but sensitive information still needs protection as it moves between those environments. Protegrity brings data-centric protection into Google Cloud so selected sensitive fields can be tokenized or encrypted before they are stored, analyzed, shared, or used by downstream workloads.
With Cloud Protect, protection operations can run through cloud-native services inside the Google Cloud environment. BigQuery workloads can incorporate protection and controlled unprotection through remote-function patterns, while other applications and pipelines can connect to Protegrity protection capabilities through supported APIs.
Key Integration Feature
Protegrity brings policy-based data protection into Google Cloud processing paths instead of requiring sensitive information to be routed to a separate protection environment. This allows organizations to apply tokenization and other protection methods close to where cloud data is processed.
For BigQuery, remote functions provide a Google Cloud-native way to invoke external protection logic directly from GoogleSQL workflows. Protegrity Cloud Protect supplies the protection service while Protegrity policy governs how selected data is protected and when clear values may be returned.
Features & Capabilities
Protect sensitive data within Google Cloud while keeping it useful for analytics, applications, and downstream processing.
01
Vaultless Tokenization for Cloud Analytics
Why It Matters
Replace sensitive data with format-preserving tokens that maintain analytical utility within BigQuery and Vertex AI, eliminating the latency of traditional lookup tables. This ensures protected data retains its original format (e.g., date or credit card length), allowing downstream GCP applications to function without breaking.
How It works
A major retailer stores tokenized customer behavior data in BigQuery; because the tokens pass standard validation checks, their AI team can train recommendation models in Vertex AI using production-volume data without ever exposing real customer identities.
02
High-Velocity Serverless Protection (via Google Cloud Functions)
Why It Matters
Leverage the massive parallelism of Google Cloud Functions to auto-scale protection and de-protection operations instantly. This eliminates bottlenecks during high-volume ETL jobs or real-time streaming, as the serverless architecture spins up thousands of concurrent instances to match your workload.
How It works
A healthcare analytics firm processes nightly patient records via Cloud Dataflow; by offloading protection to Protegrity-powered Cloud Functions, they reduced batch encryption time from 4 hours to under 20 minutes due to massive parallel throughput.
03
User-Aware Policy Enforcement
Why It Matters
Enforce strict “need-to-know” access principles by dynamically resolving the Google IAM identity of the user executing the query. This allows a single BigQuery dataset to safely serve multiple departments, with data visibility adjusting automatically in real-time based on the user’s role.
How It works
When querying a unified employee table in BigQuery, a Regional Manager sees full names and salaries, while a Data Scientist querying the exact same view sees only tokenized IDs and masked salary ranges for trend analysis.
04
IAM-Integrated Secure Authentication
Why It Matters
Enhance security posture and developer productivity by utilizing Google Service Accounts and OIDC authentication for API calls. This removes the need to hardcode API keys or manage manual secrets within your SQL scripts, Cloud Run services, or notebooks.
How It works
Data engineers utilize a “BigQuery Remote Function” connection where Google automatically handles the authentication handshake with the Protegrity Cloud Function, preventing accidental exposure of credentials in source code or logs.
05
Tag-Based Policy Automation (DataPlex & BigQuery)
Why It Matters
Scale security governance efficiently by attaching Protegrity policies to metadata tags (e.g., DataPlex tags or BigQuery Policy Tags) rather than individual columns. This ensures consistent protection is instantly applied across thousands of tables across your data estate.
How It works
A financial institution tags all social security numbers across 500 tables as “Sensitive-PII”; a single rule instantly enforces masking on every tagged column for any user who does not possess the specific “Compliance-Officer” IAM role.
Architecture &
Sample Data Flow
Protegrity’s architecture is built on a serverless design that integrates deeply with Google Cloud’s native components. By deploying protection via Google Cloud Functions, Protegrity establishes a centralized security layer that connects seamlessly with BigQuery Remote Functions. This architecture allows you to extend Protegrity’s granular protection policies across your entire GCP estate, ensuring consistent security whether data is being processed in SQL analytics, streaming data pipelines, or AI model training
The data journey
Visualizing the data journey
The data journey
The data journey explained
- 01
Ingest and Identify Sensitive Data
Data enters Google Cloud from applications, files, databases, streams, or other upstream systems and moves into the organization’s cloud processing environment.
Protection can be positioned according to where the organization wants sensitive values to transition into a protected state. - 02
Protect Sensitive Values
Configured sensitive fields are passed through Protegrity protection before continuing through downstream cloud workflows.
Depending on policy and the protection method selected, values can be tokenized or encrypted so downstream services operate on protected representations rather than unnecessarily exposing the original data. - 03
Process Protected Data
Protected information can continue into BigQuery and other supported Google Cloud processing environments.
Where the selected protection method retains the characteristics required by the workflow, protected values can continue supporting analytics, reporting, data movement, and other downstream processing without first being returned to clear form. Protegrity specifically positions protected BigQuery data for continued analytics use. - 04
Unprotect by Policy
When an approved workflow requires access to the original value, controlled unprotection can be invoked at the appropriate point in the processing path.
Protegrity policy determines how the protection operation is handled, allowing sensitive data to remain protected until clear values are required for an authorized use.
Use Cases
See how Protegrity can help organizations protect sensitive data across Google Cloud analytics, cloud migration, and downstream AI workloads while maintaining centrally managed protection policy.
AI & GenA
Protect Sensitive Data Across the AI Lifecycle
Challenge
Organizations using Google Cloud AI services need access to high-value enterprise data for model training, analytics, retrieval, and agentic workflows. Much of that data may contain PII, PHI, PCI, or other sensitive information that should not be broadly exposed as it moves from BigQuery into AI and machine learning workflows. As AI systems use more enterprise context, organizations also need a consistent way to control how sensitive information is used across model inputs, retrieved data, generated responses, and downstream analytics.
Solution
Protegrity applies data-centric protection at multiple points across the Google Cloud AI workflow. Sensitive data can be tokenized, anonymized, or otherwise protected before it is used for model training or downstream processing, allowing protected datasets in BigQuery to support analytics and machine learning while reducing exposure of original identities. For generative AI and agentic workflows, protection and policy controls can also be applied to sensitive information used in prompts, retrieved context, and generated responses. Where authorized users or applications require original values, controlled unprotection can occur according to Protegrity policy and identity.
BigQuery Remote Functions provide an additional integration point for applying protection and controlled unprotection within SQL-based analytics workflows that may support downstream AI use cases.
Result
Organizations can make sensitive enterprise data more usable across Google Cloud analytics and AI workflows while keeping protection active throughout more of the data lifecycle. Protected data can support model training, analytics, retrieval, and other downstream processing without requiring sensitive values to remain broadly available in clear form, while authorized access to original data remains governed by policy.
Healthcare Payers
Protect Sensitive Data for Cloud Analytics
Challenge
Healthcare organizations moving regulated data into Google Cloud need to make that information available for analytics without unnecessarily distributing clear patient identifiers across data pipelines, warehouses, and downstream services.
As data moves into BigQuery and other cloud processing environments, teams also need a consistent way to control which users and services can protect or re-identify sensitive information.
Solution
Protegrity Cloud Protect can apply field-level tokenization or encryption as sensitive data enters Google Cloud processing workflows. Within BigQuery, Protegrity integrates through Remote Functions that invoke the cloud-native protection service, while Enterprise Security Administrator policy governs protect and unprotect privileges.
Protected values can remain available for analytics in BigQuery, and Protegrity’s documented tokenization capabilities support format- and length-preserving options as well as join-preserving tokens.
Result
Healthcare teams can make protected datasets available for approved analytics while reducing the amount of clear sensitive information exposed across the cloud environment. Authorized users or service accounts can re-identify selected data when their Protegrity policy privileges permit it.
DEPLOYMENT
Protegrity supports Google Cloud deployments that bring field-level protection into BigQuery and broader cloud data-processing workflows while keeping protection policy centrally managed. The deployment model can combine cloud-native protection services, BigQuery integration, API-based protection, policy distribution, and audit logging based on where sensitive data is processed and how it needs to be accessed.
BigQuery Protector
Cloud API
Centralized Policy
Audit Logging
Deployment Considerations
Protegrity exposes configuration for instance limits, memory, CPU, concurrency, and timeout settings, so production workloads should be tested and sized according to the organization’s processing requirements rather than relying on a universal performance claim.
RESOURCES
Explore Protegrity documentation and technical resources for deploying, configuring, and managing data protection across Google Cloud and BigQuery environments.
Protegrity Documentation Center
Access technical documentation for Protegrity products, data protection methods, protectors, policy management, deployment, configuration, and platform capabilities.
READ MOREBigQuery Protector Documentation
Learn how Protegrity Cloud Protect integrates with BigQuery to support field-level protection, Remote Functions, policy-controlled access, and analytics on protected data.
READ MOREFrequently
Asked Questions
Protegrity supports a wide range of GCP services including BigQuery (via Remote Functions), Vertex AI, Cloud Dataflow, Cloud Run, and Google Cloud Storage. The integration is primarily architected around Google Cloud Functions, allowing any service capable of making an API call to leverage high-speed protection.
Protegrity supports on-premise, hybrid, and multi-cloud environments both within and beyond GCP. You can deploy the platform on-premise using Gateways or securely connect on-premise applications to Protegrity Cloud Protect running on GCP. For data protected in BigQuery but moved elsewhere (e.g., to an on-prem AWS environment), the same centralized policies ensure data can be seamlessly unprotected using Protegrity’s interoperable protectors.
Protegrity supports vaultless tokenization, encryption, masking, hashing, and format-preserving encryption, all centrally managed in the Enterprise Security Administrator (ESA). This centralized approach ensures that policies (e.g., “Mask Credit Cards”) are defined once and enforced consistently across BigQuery, Vertex AI, and non-GCP systems. Policy management, key rotation, and separation of duties are all handled natively
Our customers benefit from:
- Seamless BigQuery Integration: Run secure SQL analytics on protected data using BigQuery Remote Functions.
- Serverless Scalability: Support scalable protection operations through Google Cloud Functions infrastructure.
- Secure AI Workflows: Protect sensitive data before it is used in downstream Vertex AI and Gemini workflows.
- Consistent Governance: Apply centralized protection policy across Google Cloud and connected data environments.
Protegrity integrates with BigQuery via Remote Functions (SQL UDFs that call Cloud Functions) and with Vertex AI via secure API calls within data pipelines. This allows users to query data or train models using standard Google tools while Protegrity handles the cryptographic operations in the background. Please refer to the Integration Features section for more details.
See the Protegrity
platform in action
Accelerate data access and turn data security into a competitive advantage with Protegrity’s uniquely data-centric approach to data protection.
Get an online or custom live demo.