Plain-English Explanation
What this episode is about
This episode discusses how law firms manage the huge amount of information spread across document-management systems, cloud drives, email, collaboration apps, archives, and artificial-intelligence tools.
Tony Forde argues that firms need to:
•Find and classify information wherever it is stored.
•Remove unnecessary or outdated files.
•Keep reliable records of who approved what and when.
•Let attorneys work in the software they already prefer.
•Carefully prepare data before using AI.
•Start with a small pilot project instead of trying to reorganize the entire firm at once.
The central message is: good information governance is becoming the foundation for trustworthy AI.
Main ideas in simple terms
1. Important information is no longer in one central system
Law firms traditionally treated a document management system as the official home for work product. But attorneys now work in many places:
•Microsoft SharePoint and OneDrive
•Box and ShareFile
•Email
•Teams and other collaboration platforms
•E-discovery software
•Personal or departmental file shares
•AI tools
As a result, the firm’s official system may contain only part of the information that matters.
Tony compares this to having the firm’s records scattered across many rooms, with no reliable inventory of what is in each room.
2. Cloud storage creates a “hidden attic”
Because cloud storage is easy and relatively cheap, people often keep everything:
•Old drafts
•Duplicate files
•Personal working copies
•Completed matters
•Files nobody has opened for years
•Client information that belongs in the official system
A firm may discover that hundreds of terabytes are sitting in OneDrive, SharePoint, or Box. Much of it may be unnecessary, misplaced, or legally sensitive.
This unmanaged information is often called dark data.
3. ROT must be identified and removed
ROT means:
•Redundant: duplicate copies
•Obsolete: no longer useful or current
•Trivial: low-value material, such as temporary notes or routine clutter
The firm must decide what should be:
•Kept in the official document system
•Archived
•Returned to the client
•Deleted or destroyed
This process is called disposition. The goal is not simply to delete files, but to make the decision according to documented rules and approvals.
4. Different generations of attorneys work differently
Tony says a law firm may effectively contain several technology cultures at once.
Some senior lawyers may prefer paper, folders, and simple tools. Younger lawyers may expect cloud systems, mobile access, collaboration platforms, and AI.
A governance program will fail if it forces everyone into an unfamiliar workflow. The better approach is to connect to the places where attorneys already work and provide governance functions there.
5. AI needs clean, well-organized data
AI systems learn from the information they can access. If that information is duplicated, incomplete, mislabeled, or outdated, the AI may produce unreliable results.
Before using AI effectively, firms may need to connect information from several sources, such as:
•The document-management system
•Time-billing records
•Employee directories
•Matter and client records
•Collaboration systems
This helps the system understand not only what a document says, but also:
•Who worked on it
•Which client or matter it belongs to
•When it was created or changed
•Whether it is authoritative or merely a draft
6. Humans still need to supervise AI classification
Tony warns against assuming that AI can automatically organize every document correctly.
AI may make two kinds of mistakes:
•False positive: it labels something as belonging to a category when it does not.
•False negative: it fails to recognize something that does belong in that category.
Simple data, such as phone numbers or credit-card numbers, is relatively easy for software to identify. More complex legal content is much harder. A contract, intellectual-property document, or legal argument may require genuine understanding of context.
Tony therefore favors a system in which people first provide examples and guidance, and the machine then assists with the work.
7. Governance should fit into existing attorney workflows
Attorneys may prefer to review information in Relativity, Everlaw, SharePoint, or the document-management system.
Tony’s company tries to deliver governance tasks inside those tools where possible. This reduces training and makes adoption easier.
The principle is simple: bring governance to the user instead of expecting every user to change habits.
8. Audit trails make decisions defensible
A firm must be able to show:
•What information was reviewed
•Who reviewed it
•What decision was made
•Who approved the decision
•When each step occurred
•Whether anything was changed or transferred
This record is called an audit trail or audit log.
It is important because a firm may later need to prove to a regulator, auditor, court, client, or opposing party that it handled information responsibly.
9. Approval workflows should be controlled
Some decisions can be handled by several people at the same time. Others must follow a specific order.
For example:
1. A secretary identifies old client files.
2. A practice-group leader reviews them.
3. Records management confirms the retention rule.
4. A responsible partner gives final approval.
5. The files are archived, transferred, or destroyed.
A system should record every step. Otherwise, someone may later discover that a file was destroyed or transferred with only one person’s informal approval.
10. AI can create a new data problem
AI does not only consume data; it also creates more of it:
•Prompts
•AI-generated answers
•Uploaded documents
•Downloaded results
•Conversation histories
•Model-evaluation records
If these materials are not governed, firms may replace one information mess with another.
Tony recommends controlling prompts and outputs within the firm’s own environment where practical, while monitoring usage and costs.
11. AI usage can become unexpectedly expensive
Some AI services charge according to usage. Large document-processing projects can create large bills, especially when many files are uploaded or analyzed at once.
Tony describes a project involving roughly 60 terabytes of data and emphasizes that the computing power required can cause significant cost spikes.
Firms therefore need:
•Usage monitoring
•Spending limits
•Clear ownership of AI environments
•Rules about what data may be uploaded
•Records of prompts and outputs
12. Start with a small, measurable project
Tony recommends beginning with one matter, department, or practice group.
A pilot project might answer questions such as:
•What information is stored there?
•How much is duplicated?
•What has not been touched in years?
•Which files should be retained?
•Which files can be archived or destroyed?
•How long does the approval process take?
A successful small project creates evidence and supporters inside the firm. Those supporters can then help expand the program.
Technical terms explained
•AI (Artificial Intelligence): Computer software that performs tasks associated with human intelligence, such as finding patterns, summarizing text, or answering questions.
•AI governance: Rules and controls for using AI safely, consistently, legally, and affordably.
•API (Application Programming Interface): A structured way for two software systems to exchange information automatically.
•Authoritative source: The official, trusted location for a particular type of information. In this discussion, the document-management system is intended to be the official record.
•Auto-classification: Software automatically assigning documents to categories, such as client matter, contract, invoice, or personnel record.
•Batching: Processing information in groups instead of handling each item individually. This can make classification more consistent and easier to review.
•Cloud storage: Data stored on computers operated by a service provider and accessed over the internet, rather than stored only on the firm’s own machines.
•Classification: Assigning information labels or categories so it can be searched, protected, retained, or deleted appropriately.
•Collaboration platform: Software that allows people to share, edit, discuss, and manage files together, such as Microsoft Teams or SharePoint.
•Comparative source: Another source of information used to check or enrich the meaning of a document—for example, comparing a document with billing records, a directory, or a matter database.
•Custodian: The person who possesses, controls, or is responsible for potentially relevant information, especially during litigation or investigation.
•Data governance: The broader management of data’s quality, ownership, security, access, and lifecycle.
•Data minimization: Collecting, storing, or using only the information actually needed for a legitimate purpose.
•Dark data: Information an organization stores but does not properly understand, classify, use, or manage.
•Defensible disposition: Deleting, archiving, transferring, or returning information according to documented rules, consistently applied procedures, and proper approvals.
•DMS (Document Management System): A system designed to organize official documents, often by client, legal matter, document type, version, and permissions. Examples include iManage and similar products.
•E-discovery: The process of finding, collecting, reviewing, and producing electronically stored information for a legal case or investigation.
•Endpoint: A place where information is stored or accessed, such as OneDrive, SharePoint, email, Box, a file share, or an AI application.
•False negative: A failure to identify something that should have been identified—for example, failing to classify a confidential client document correctly.
•False positive: Incorrectly identifying something as belonging to a category—for example, labeling an ordinary memo as a privileged legal document.
•Fine-grained retention rule: A detailed rule specifying exactly how long a particular type of information must be kept and what should happen afterward.
•Integration: A connection that allows different software systems to exchange data or trigger actions.
•ISO certification: Certification showing that an organization follows a recognized international management standard. In this context, it may relate to information security or process controls.
•Legal hold: An instruction to preserve potentially relevant information because litigation, an investigation, or a legal dispute is expected or already underway.
•Machine learning: A form of AI in which software learns patterns from examples or data instead of being programmed with every rule manually.
•Matter: A specific legal engagement or case handled for a client.
•OneDrive: Microsoft’s cloud file-storage service, often used for individual work or sharing.
•Permissioning: Controlling who can view, edit, download, share, or delete information.
•PII (Personally Identifiable Information): Information that can identify a person, such as a name, address, phone number, Social Security number, or account number.
•Prompt: The instruction or question given to an AI system.
•Prompt library: A controlled collection of approved, reusable prompts for common tasks.
•Retention: Keeping information for a defined period because it is useful, legally required, or part of the organization’s official records.
•ROT: Redundant, Obsolete, and Trivial information.
•RMS (Records Management System): Software that manages official records, including retention periods, classification, holds, and final disposition.
•Sandbox: A separated, controlled environment where software or experiments can be used without exposing the wider organization to unnecessary risk.
•SharePoint: Microsoft’s platform for storing, sharing, and collaborating on documents and other workplace information.
•SOC 2: An independent auditing framework that evaluates whether a service provider has appropriate controls for security, availability, confidentiality, privacy, and processing integrity.
•Source of truth: The location an organization treats as the most reliable and official version of information.
•Structured data: Information arranged in predictable fields, such as names, dates, matter numbers, or billing entries.
•Unstructured data: Information without a consistent database format, such as emails, documents, chat messages, and notes.
•Vector: In AI, a numerical representation of the meaning or characteristics of text or other data. Vectors help systems find information that is conceptually similar, even when the exact words differ.
•Workflow: A defined sequence of tasks, decisions, and approvals.
Why this matters
Law firms are becoming dependent on information spread across many systems. That creates practical, legal, security, and financial risks:
•Important documents may be lost or overlooked.
•Confidential information may be stored in the wrong place.
•Old files may remain accessible indefinitely.
•The firm may be unable to prove why it deleted or retained something.
•AI may learn from inaccurate or inappropriate material.
•AI-generated data may create a second wave of unmanaged information.
•Usage-based AI costs may become difficult to control.
The episode’s practical advice is straightforward: inventory the information, clean up ROT, establish clear approval and retention rules, connect governance to existing tools, and begin with a small project that produces measurable results.
A well-governed information environment does more than reduce clutter. It gives attorneys safer access to the right material and gives AI a cleaner, more trustworthy foundation.