The modern art of ediscovery data filtering
eDiscovery culling and filtering techniques to defensibly reduce document volumes — and cut review costs before they spiral


According to eDiscovery expert Michael Arkfeld, up to 98% of data collected in response to an ediscovery request turns out to be non-responsive. The attorneys who know how to filter strategically before review begins control the scope, the timeline, and the cost of any matter. Those who don't end up reviewing mountains of irrelevant data at full price.
What is ediscovery data filtering?
Ediscovery data filtering is the process of narrowing a collected data set down to what's actually relevant to a case — using techniques like deduplication, date ranges, file type exclusions, keyword searches, and predictive coding — before documents ever reach a reviewer's screen. Done correctly, it's not just a cost-cutting step. It's a defensible, documented part of the discovery process that courts expect the attorney of record to understand and stand behind.
This guide covers the filtering and culling techniques that make the difference — from early case assessment and deduplication to keyword strategies, hash values, and predictive coding — along with the defensibility standards you need to document your work.
What is predictive coding in ediscovery?
Predictive coding, or technology-assisted review (TAR), uses machine learning to rank documents by probability of relevance. In high-volume matters, it acts as a second filter after technical filtering — helping legal teams identify and set aside a pool of likely-irrelevant documents rather than reviewing everything by hand. Litigator Ralph Losey recommends retaining only documents with a 90% or higher probability of relevance, with the exact threshold negotiated between parties.
How do you validate a defensible data filtering process?
Courts expect documented proof, not just a clean result. A defensible process means keeping a chain-of-custody log covering every step from collection through production, along with testimony from whoever collected the ESI about how it was handled. Your ediscovery vendor should also provide four reports: an extraction report (what was pulled from the source data), a processing report (what was ingested and exported), a de-duplication report (what duplicates were removed and which file is the master), and an exceptions report (what failed to process and why).
What happens if ediscovery data is lost or destroyed?
Federal Rule of Civil Procedure 37(e) replaced the old "safe harbor" provision. If lost electronically stored information can be replaced or restored, no sanction applies. If it can't be replaced and the requesting party is prejudiced, courts can order sanctions "no greater than necessary to cure that prejudice" — which is why documenting what was collected, what was discarded, and why is a legal safeguard, not just good hygiene.
How much can data filtering reduce ediscovery costs?
Since up to 98% of collected data is typically non-responsive, filtering before review is where the real savings happen. Every document culled before it reaches a reviewer is a document you never pay review rates on — which is why a strong filtering strategy, not a bigger review team, is usually the fastest way to bring a runaway budget back under control.
5 things you'll learn:
- Why eDiscovery data filtering is a legal obligation, not just a cost-saving move — and what Rule 26(g) requires of the attorney signing off
- How to use early case assessment, deNISTing, deduplication, and email threading to cut your document set before review begins
- How to build a defensible filtering strategy using date ranges, file types, custodians, and keyword searches
- What hash values are, why they matter for chain of custody, and how they underpin most filtering technology
- When to consider predictive coding and technology-assisted review for high-volume matters
What's inside:
- A technical overview of the full data filtering and culling workflow, from raw ESI to production-ready sets
- Filtering strategy frameworks for both producing and requesting parties
- Guidance on validating processed data and maintaining a defensible chain of custody
- A diagram of the complete technical filtering process — deNIST, dedup, date range, domain, file type, custodian, and email threading
Download the guide and go into your next matter with a smarter, leaner approach to data.
Grab the guide
Real-world insight from the Nextpoint services team, delivered straight to your inbox. Whether you're tackling a complex investigation or streamlining day-to-day discovery, our library of practical guides, checklists, and on-demand webinars gives you something useful at every stage of the case lifecycle.
Related Resources
Experience Nextpoint for yourself
Learn how our transparent pricing and powerful platform help legal teams streamline litigation from discovery to decision.



