Why would you use Purview Data Loss Prevention with Copilot?

One of the questions I (still) get is:

"We're interested in Microsoft 365 Copilot, but we're worried that staff might discover information they shouldn't have access to."

It's a reasonable concern, even though Copilot has been out for a few years :)

The first response to the question is that Copilot respects existing permissions. If a user does not have access to a file, Copilot cannot magically access it on their behalf.

But there is a second part to the conversation because some organizations cannot rely solely on permissions. In the real world:

  • Permissions are often messy.

  • Sensitive information is frequently overshared.

  • Legacy content may have been accessible for years, but not easy to find. i.e. "Security by obscurity"

  • Users sometimes paste confidential information directly into AI prompts.

  • Employees may use both Microsoft Copilot and third-party AI tools.

This is where Microsoft Purview Data Loss Prevention (DLP) becomes an important part of a Copilot governance strategy.

While permissions determine what users can access, DLP helps control how sensitive information can be used within AI experiences. Microsoft has added several DLP capabilities for Microsoft 365 Copilot and Copilot Chat that allow organizations to reduce risk without blocking productivity.

Think of DLP as a second layer of protection

A simple way to explain the relationship is:

Control Purpose
Permissions Controls what users can access
Sensitivity Labels Identifies important content
DLP Controls how sensitive data can be used
Records Management Controls how long content should exist

Organizations need all four, particularly when introducing AI.

Permissions remain the foundation, and it's still CRITICAL to invest in proper information management, but DLP can help address situations where sensitive information is technically accessible but should not be used in certain AI scenarios.

Scenario 1: Stop users from pasting sensitive information into Copilot

One of the simplest and most effective controls is prompt protection.

Consider a user who copies information from a spreadsheet containing:

  • Payroll information

  • Social Insurance Numbers

  • Credit card information

  • Customer financial data

and pastes it directly into a Copilot prompt.

Microsoft Purview DLP can identify sensitive information types and prevent Copilot from processing the request.

Example of pasting sensitive data into Copilot (Image source: Microsoft)

Where to start

Don't try to build hundreds of custom rules immediately.

Start with Microsoft's built-in sensitive information types such as credit card numbers, banking info, passport numbers etc. These already cover many common risk scenarios.

Scenario 2: Keep Copilot away from your most sensitive content

Many organizations have information that deserves additional protection, including:

  • Executive discussions

  • Legal investigations

  • Mergers and acquisitions content

  • HR investigations

  • Board materials

  • Intellectual property such as patented formulas

  • etc.

Purview DLP can prevent Copilot from using files that have specific sensitivity labels when generating responses.

Where to start

Instead of attempting to classify everything, start by identifying a small set of highly sensitive labels such as:

  • Highly Confidential

  • Executive Confidential

  • Legal Privileged

These categories often provide the biggest risk reduction with the least effort.

Then use DLP to restrict content with those labels from being used in Copilot to generate summaries and responses.

Scenario 3: Reduce use of external and untrusted content

Organizations are increasingly concerned about users relying on information from external sources, including email and AI applications. Why? In order to mitigate the risk of relying on untrusted data sources, or the risk of prompt injection.

Microsoft has introduced controls that can prevent external email content from being used as grounding data for Copilot responses (currently in preview).

When this control is enabled, Copilot excludes emails from external domains from being referenced or summarized during prompt processing, while continuing to use internal Microsoft 365 data sources where permitted.

Where to start

Evaluate whether specific business groups should rely primarily on trusted internal knowledge when using Copilot for decision-making activities. Make a plan to simulate then deploy the policy where needed.

Scenario 4: Limit sensitive data sent to web grounded prompts

Copilot is useful because it uses both web-grounding and internal sources to generate responses.

However, organizations often do not want sensitive information included in scenarios involving external web searches.

Microsoft now allows organizations to detect sensitive information in prompts and prevent web grounding from being used. Copilot can still generate responses using approved Microsoft 365 content without involving external web sources.

Where to start

This is often a low-friction policy to implement because it has minimal impact on normal business use while helping address a common leadership concern around data leakage.

Scenario 5: Protect against third-party AI risks

Sometimes in organizations, the larger concern is staff using public AI services and pasting sensitive organizational information into them.

Several Purview capabilities can help monitor and protect against sensitive information being shared with external AI services.

Where to start

Focus first on protecting:

  • Employee information

  • Customer information

  • Financial information

  • Intellectual property

These categories usually represent the highest business risk.

Final thoughts

A common misconception is that Copilot security is only about permissions. Permissions are critical, but they are not sufficient on their own.

I think of it this way:

Permissions determine what users can access. DLP helps determine how sensitive information can be used once they have access to it.

To get started with DLP, don't try to implement dozens of policies before understanding your organization's data. A better approach is to start with highest-risk scenarios or sensitive data, then monitor and improve.

Learn more

Next
Next

Long reads that are impossible to put down