Back to blog Security

Privilege and Confidentiality in AI-Assisted Due Diligence

Amara Okonkwo
Privilege and confidentiality in AI document review

Every in-house legal team evaluating AI-assisted contract review eventually arrives at the same question: what happens to the documents we upload? The question is not paranoid. Purchase agreements contain deal terms that are confidential between the parties. Attorney-reviewed drafts may contain work product. Disclosure schedules reference the target company's customers, financial data, and pending legal matters. Uploading these documents to a third-party service creates legitimate questions about data handling, training practices, and whether the confidential nature of the materials is preserved on the other side of the API call.

The right answer to that question is not "AI tools are risky, so don't use them." That posture treats all tools as equivalent when they differ materially in how they handle client data. The right answer starts with asking the right questions about the specific tool under evaluation.

The Privilege Question

Attorney-client privilege protects communications between a lawyer and client made in confidence for the purpose of legal advice. Work product doctrine protects materials prepared in anticipation of litigation. For either protection to apply, the communication or document must have been kept confidential.

When a legal team uploads a purchase agreement to a third-party AI service, the threshold privilege question is whether that act constitutes a voluntary disclosure to a third party. Voluntary disclosure to a third party not necessary to facilitate the legal representation can waive attorney-client privilege. The analysis of whether uploading to an AI review tool waives privilege depends on the specific facts.

A purchase agreement itself is typically not a privileged document. It is a contract between commercial parties. The act of uploading it to an AI tool for clause extraction is unlikely to implicate privilege in most cases. But the same upload act applied to a document that contains attorney mental impressions, a negotiation strategy memo attached as an exhibit, or a redline with attorney comments does raise genuine privilege considerations. The relevant question is not whether AI tools in general create privilege exposure, but whether the specific document being uploaded and the specific service being used create that exposure for that document.

The more common concern in practice is not technical privilege waiver but practical confidentiality. Purchase agreements contain commercially sensitive information that neither party wants shared beyond the deal team. Disclosure schedules contain customer lists, financial data, and pending litigation details that are highly sensitive regardless of privilege status. For these documents, the practical question is data handling rather than privilege doctrine.

Model Training: The Right Question to Ask

The most significant data handling concern for legal teams using AI contract review tools is whether uploaded documents are used to train the service's underlying models. If they are, the target company's confidential contract terms, customer relationships described in the disclosure schedules, and financial data provided to the AI become part of the training corpus for a model that will subsequently respond to queries from other users.

This is not a hypothetical risk. Several general-purpose AI services that were used for document-related tasks in their early deployment used user data as training input under terms of service that permitted it. The data was nominally anonymized, but the practical risk for a purchase agreement containing specific contract terms and deal economics is that those terms could surface in responses to other users with sufficient prompting.

The right question when evaluating an AI contract review tool is specific: Does the service use documents uploaded for processing as training data for its models? If yes, under what circumstances, and with what opt-out mechanism? If the answer is ambiguous or the terms of service are unclear, that ambiguity itself is an answer. A tool designed for professional legal use should have a clear and explicit policy that client documents are not used for model training.

For Statuteharbor, the policy is explicit: documents uploaded for review are not used to train models. The answer to the training question should not require reading through layered terms of service or relying on an implied understanding. It should be a stated policy that legal teams can point to in the event of any client or regulatory inquiry about how their documents were handled.

Data Residency and Retention

Beyond the training question, legal teams should understand where their data is stored, for how long, and who can access it. For most in-house legal teams, the primary concern is ensuring that the vendor's data practices are consistent with any confidentiality obligations they have to the parties to the transaction.

Data residency matters because documents containing deal terms may be subject to data localization requirements in some jurisdictions. A cross-border acquisition where one party is subject to GDPR data handling requirements, for example, may require that data related to EU-based entities is processed and stored within the EU. Using a US-based AI review service without understanding its data residency may create compliance exposure in those situations.

Retention policy matters because the longer a service retains uploaded documents, the longer the confidentiality exposure window remains open. A service that retains uploaded documents indefinitely creates ongoing exposure that a service with a clear document deletion policy after a defined period does not.

Access controls are the third element. Who within the vendor's organization can access uploaded documents, under what circumstances, and with what logging? For a tool used to review M&A purchase agreements, the answer should be: access is limited to infrastructure teams for maintenance purposes, access events are logged, and human review of client documents does not occur except for identified support cases with client authorization.

What the Contract with the Vendor Should Say

The best documentary protection for these questions is a clear vendor contract that addresses them directly. A data processing agreement or similar instrument should cover: the specific purposes for which the vendor may process uploaded data, whether model training is a permitted use, data retention and deletion schedules, subprocessor disclosure, and the vendor's obligations on breach notification.

Legal teams that are evaluating AI contract review tools should not rely on the vendor's public-facing statements about data handling without a contractual commitment to back them. Public statements can change. Contractual obligations create enforceable commitments. For tools that will be used to review sensitive deal documents, the evaluation process should include reviewing and potentially negotiating the data processing terms before procurement.

This is not unique to AI tools. Any cloud-based document service used for M&A work should be subject to the same due diligence. What makes AI contract review tools different is that the training-data question is specific to this category of service and may not be covered by template data processing agreements that were written before AI training practices became a common concern.

Privilege Logs and AI-Assisted Review

One privilege-adjacent question that arises when AI tools are used in the context of litigation support or regulatory investigation is whether the fact of AI assistance needs to be disclosed in a privilege log. This question is currently unsettled and varies by jurisdiction and context.

For M&A transaction review, however, the privilege log question is generally not the central issue because the documents being reviewed are typically not privileged. Purchase agreements, disclosure schedules, and supporting due diligence documents are business records, not attorney communications. The privilege considerations at play in M&A document review are the practical confidentiality concerns about deal terms rather than the technical privilege log questions that arise in litigation contexts.

Where privilege does arise in M&A review is in the attorney's analysis layer: the written summary of issues, the negotiation strategy memo, and the advice provided to the client based on the review. Those documents do not need to pass through the AI tool. The AI tool extracts the clause terms from the agreement; the attorney applies judgment to those terms in communications that remain within the attorney-client relationship and are not uploaded to any external service.

The Practical Standard

The practical standard for evaluating AI contract review tools from a confidentiality perspective is not perfection. No data handling arrangement eliminates all risk. The standard is whether the tool's data practices create material additional risk relative to other document handling practices already in use.

Legal teams that use cloud-based document management systems, share deal documents through virtual data rooms hosted by third-party providers, and transmit purchase agreements via standard email have already made data handling decisions with comparable risk profiles. The question for AI review tools is whether the additional data handling they require is within the risk parameters those teams have already accepted for other parts of the deal workflow, with the specific training-data question addressed explicitly by the vendor's policy. If the answer to that question is yes, the confidentiality analysis supports adoption. If the training-data policy is unclear or the vendor cannot provide a clear data processing commitment, the analysis does not support it.

Try it yourself

Review your next purchase agreement with Statuteharbor

Upload a PDF and get a structured flag report in minutes. No credit card required for the 14-day trial.