{"id":2021,"date":"2014-03-14T16:41:08","date_gmt":"2014-03-14T16:41:08","guid":{"rendered":"http:\/\/jolt.richmond.edu\/?p=2021"},"modified":"2019-03-08T19:52:30","modified_gmt":"2019-03-09T00:52:30","slug":"defensible-data-deletion-a-practical-approach-to-reducing-cost-and-managing-risk-associated-with-expanding-enterprise-data-2","status":"publish","type":"post","link":"https:\/\/blog.richmond.edu\/jolt\/2014\/03\/14\/defensible-data-deletion-a-practical-approach-to-reducing-cost-and-managing-risk-associated-with-expanding-enterprise-data-2\/","title":{"rendered":"Defensible Data Deletion: A Practical Approach to Reducing Cost and Managing Risk Associated with Expanding Enterprise Data"},"content":{"rendered":"<p style=\"text-align: left\" align=\"center\"><a href=\"http:\/\/jolt.richmond.edu\/v20i2\/article6.pdf\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-128\" alt=\"pdf_icon\" src=\"http:\/\/jolt.richmond.edu\/files\/2012\/05\/pdf_icon1.gif\" width=\"16\" height=\"16\" \/>DownloadPDF<\/a><\/p>\n<p style=\"text-align: center\" align=\"center\">Cite as: Dennis R. Kiker,\u00a0<i>Defensible Data Deletion: A Practical Approach to Reducing Cost and Managing Risk Associated with Expanding Enterprise Data<\/i>, 20 Rich. J.L. &amp; Tech. 6 (2014), http:\/\/jolt.richmond.edu\/v20i2\/article6.pdf.<\/p>\n<p style=\"text-align: center\" align=\"center\"><b>\u00a0<\/b><\/p>\n<p style=\"text-align: center\">Dennis R. Kiker*<i><\/i><\/p>\n<p align=\"center\"><b style=\"line-height: 1.714285714;font-size: 1rem\">\u00a0<\/b><\/p>\n<h2 align=\"center\"><b>I.\u00a0 Introduction<\/b><\/h2>\n<p style=\"text-align: left\" align=\"center\">\n<p style=\"text-align: left\" align=\"center\">[1]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 Modern businesses are hosts to steadily increasing volumes of data, creating significant cost and risk while potentially compromising the current and future performance and stability of the information systems in which the data reside.\u00a0 To mitigate these costs and risks, many companies are considering initiatives to identify and eliminate information that is not needed for any business or legal purpose (a process referred to herein as \u201cdata remediation\u201d).\u00a0 There are several challenges for any such initiative, the most significant of which may be the fear that information subject to a legal preservation obligation might be destroyed.\u00a0 Given the volumes of information and the practical limitations of search technology, it is simply impossible to eliminate all risk that such information might be overlooked during the identification or remediation process.\u00a0 However, the law does not require that corporations eliminate \u201call risk.\u201d\u00a0 The law requires that corporations act reasonably and in good faith,[1] and it is entirely possible to design and execute a data remediation program that demonstrates both.\u00a0\u00a0 Moreover, executing a reasonable data remediation program yields more than just economic and operational benefits.\u00a0 Eliminating information that has no legal or business value enables more effective and efficient identification, preservation, and production of information requested in discovery.[2]<\/p>\n<p style=\"text-align: left\">\u00a0[2]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 This Article will review the legal requirements governing data preservation in the litigation context, and will demonstrate that a company can conduct data remediation programs while complying with those legal requirements.\u00a0 First, we will examine the magnitude of the information management challenge faced by companies today.\u00a0 Then we will outline the legal principles associated with the preservation and disposition of information.\u00a0 Finally, with that background, we will propose a framework for an effective data remediation program that demonstrates reasonableness and good faith while achieving the important business objectives of lowering cost and risk.<\/p>\n<p>&nbsp;<\/p>\n<h2 align=\"center\"><b>II.\u00a0 The Problem: More Data Than We Want or Need<\/b><\/h2>\n<p><b>\u00a0<\/b>[3]<b>\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 <\/b>Companies generate an enormous amount of information in the ordinary course of business.\u00a0 More than a decade ago, researchers at the University of California at Berkeley School of Information Management and Systems undertook a study to estimate the amount of new information generated each year.[3] \u00a0Even ten years ago, the results were nearly beyond comprehension.\u00a0 The study estimated that the worldwide production of original information as of 2002 was roughly five exabytes of data, and that the storage of new information was growing at a rate of up to 30% per year.[4]\u00a0 Put in perspective, the same study estimates that five exabytes is approximately equal to all of the words ever spoken by human beings.[5]\u00a0 Regardless of the precision of the study, there is little question that the volume of information, particularly electronically stored information (\u201cESI\u201d) is enormous and growing at a frantic pace.\u00a0 Much of that information is created by and resides in the computer and storage systems of companies.\u00a0 And the timeworn adage that \u201cstorage is cheap\u201d is simply not true when applied to large volumes of information.\u00a0 Indeed, the cost of storage can be great and come from a number of different sources.<\/p>\n<p>[4]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 First, there is the cost of the storage media and infrastructure itself, as well as the personnel required to maintain them.\u00a0 Analysts estimate the total cost to store one petabyte of information to be almost five million dollars per year.[6]\u00a0 The significance of these costs is even greater when one realizes that the vast majority of the storage for which companies are currently paying is not being used for any productive purpose.\u00a0 At least one survey indicates that companies could defensibly dispose of up to 70% of the electronic data currently retained.[7]<\/p>\n<p>[5]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 Second, there is a cost associated with keeping information that currently serves no productive business purpose.\u00a0 The existence of large volumes of valueless information makes it more difficult to find information that is of use.\u00a0 Numerous analysts and experts have recognized the tremendous challenge of identifying, preserving, and producing relevant information in large, unorganized data stores.[8]\u00a0 As data stores increase in size, identifying particular records relevant to a specific issue becomes progressively more challenging.\u00a0 One of the best things a company can do to improve its ability to preserve potentially relevant information, while also conserving corporate resources, is to eliminate information from its data stores that has no business value and is not subject to a current preservation obligation.<\/p>\n<p>[6]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 Eliminating information can be extremely challenging, however, due to the potential cost and complexity associated with identifying information that must be preserved to comply with existing legal obligations.\u00a0 When dealing with large volumes of information, manual, item-by-item review by humans is both impractical and ineffective.\u00a0 From the practical perspective, large volumes of information simply cannot be reviewed in a timely fashion with reasonable cost.\u00a0 For example, consider an enterprise system containing 500 million items.\u00a0 Even assuming a very aggressive review rate of 100 documents per hour, 500 million items would require five million man-hours to review.\u00a0 At any hourly rate, the cost associated with such a review would be prohibitive.<\/p>\n<p>[7]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 Even when leveraging commonly used methods of data culling to reduce the volume required for review, such as deduplication, date culling, and key word filtering, the anticipated volume would still be unwieldy when even a 90% reduction in volume would require review of 50 million items. Moreover, studies have long demonstrated that human reviewers are often quite inconsistent with respect to identifying \u201crelevant\u201d information, even when assisted by key word searches.[9]<\/p>\n<p>[8]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 Current scholarship also shows that human reviewers do not consistently apply the concept of relevance and that the overlap, or the measure of the percentage of agreement on the relevancy of a particular document between reviewers, can be extremely low.[10]\u00a0 Counter-intuitively, the result is the same even when more \u201csenior\u201d review attorneys set the \u201cgold standard\u201d for determining relevance.[11]\u00a0 Recent studies comparing technology-assisted processes with traditional human review conclude that the former can and will yield better results.\u00a0 Technology can improve both recall (the percentage of the total number of relevant documents in the general population that are retrieved through search) and precision (percentage of retrieved documents that are, in fact, relevant) than humans can achieve using traditional methods.[12]<\/p>\n<p>[9]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 There is also growing judicial acceptance of parties\u2019 use of technology to help reduce the substantial burdens and costs associated with identifying, collecting, and reviewing ESI.\u00a0 Recently, the U.S. District Court for the Southern District of New York affirmed Magistrate Judge Andrew Peck\u2019s order approving the parties\u2019 agreement to use \u201cpredictive coding,\u201d a method of using specialized software to identify potentially relevant information.<sup><sup>[13]<\/sup><\/sup><\/p>\n<p>[10]\u00a0\u00a0\u00a0\u00a0\u00a0 Likewise, a Loudon County, Virginia Circuit Court judge recently granted a defendant\u2019s motion for protective order allowing the use of predictive coding for document review.[14]\u00a0 The defendant had a data population of 250 GB of reviewable ESI comprising as many as two million documents, which, it argued, would require 20,000 man-hours to review using traditional human review.[15]\u00a0 The defendant explained that traditional methods of linear human review likely \u201cmisses on average 40% of the relevant documents, and the documents pulled by human reviewers are nearly 70% irrelevant.\u201d[16]<\/p>\n<p>[11]\u00a0\u00a0\u00a0\u00a0\u00a0 Similarly, commentary included with recent revisions to Rule 502 of the Federal Rules of Evidence indicate that using computer-assisted tools may demonstrate reasonableness in the context of privilege review: \u201cDepending on the circumstances, a party that uses advanced analytical software applications and linguistic tools in screening for privilege may be found to have taken \u2018reasonable steps\u2019 to prevent inadvertent disclosure.\u201d[17]<\/p>\n<p>[12]\u00a0\u00a0\u00a0\u00a0\u00a0 Simply put, dealing with the volume of information in most business information systems is beyond what would be humanly possible without leveraging technology.\u00a0 Because such systems contain hundreds of millions of records, companies effectively have three choices for searching for data subject to a preservation obligation: they can rely on the search capabilities of the application or native operating system, they can invest in and employ third-party technology to index and search the data in its native environment, or they can export all of the data to a third-party application for processing and analysis.<\/p>\n<p>&nbsp;<\/p>\n<h2 align=\"center\"><b>III.\u00a0 The Solution: Defensible Data Remediation<\/b><\/h2>\n<p style=\"text-align: left\" align=\"center\">[13]\u00a0\u00a0\u00a0\u00a0\u00a0 Simply adding storage and retaining the ever-increasing volume of information is not a tenable option for businesses given the cost and risk involved.\u00a0 However, there are risks associated with data disposition as well, specifically that information necessary to the business or required for legal or regulatory reasons will be destroyed.\u00a0 Thus, the first stage of a defensible data remediation program requires an understanding of the business and legal retention requirements applicable to the data in question.\u00a0 Once these are understood, it is possible to construct a remediation framework appropriate to the repository that reflects those requirements.<\/p>\n<h3>\u00a0<b>A.\u00a0 Retention and Preservation Requirements<\/b><\/h3>\n<p><b>\u00a0<\/b>[14]\u00a0\u00a0\u00a0\u00a0\u00a0 The U.S. Supreme Court has recognized that \u201c\u2018[d]ocument retention policies,\u2019 which are created in part to keep certain information from getting into the hands of others, including the Government, are common in business.\u201d[18]\u00a0 The Court noted that compliance with a valid document retention policy is not wrongful under ordinary circumstances.[19]\u00a0 Document retention policies are intended to facilitate retention of information that companies need for ongoing or historical business purposes, or as mandated by some regulatory or similar legal requirement.\u00a0 Before attempting remediation of a data repository, the company must first understand and document the applicable retention and preservation requirements.<\/p>\n<p>[15]\u00a0\u00a0\u00a0\u00a0\u00a0 It is beyond the scope of this Article to outline all of the potential business and regulatory retention requirements.[20]\u00a0 Ideally, these would be reflected in the company\u2019s record retention schedules.\u00a0 However, even when a company does not have current, up-to-date retention schedules, embarking on a data remediation exercise affords the opportunity to develop or update such schedules in the context of a specific data repository.\u00a0 Most data repositories contain limited types of data.\u00a0 For example, an order processing system would not contain engineering documents.\u00a0 Thus, a company is generally focused on a limited number of retention requirements for any given repository.\u00a0 There are exceptions to this rule, such as with e-mail systems, shared-use repositories (e.g., Microsoft SharePoint), and shared network drives.\u00a0 Even then, focusing on the specific repository will enable the company to likewise focus on some limited subset of its overall record retention requirements.\u00a0 Once a company has identified the business and regulatory retention requirements applicable to a given data repository, information in the repository that is not subject to those requirements is eligible for deletion unless it is subject to the duty to preserve evidence.<\/p>\n<p>[16]\u00a0\u00a0\u00a0\u00a0\u00a0 The modern duty to preserve derives from the common law duty to preserve evidence and is not explicitly addressed in the Federal Rules of Civil Procedure.[21]\u00a0 The duty does not arise until litigation is \u201creasonably anticipated.\u201d[22]\u00a0 Litigation is \u201creasonably anticipated\u201d when a party \u201cknows\u201d or \u201cshould have known\u201d that the evidence may be relevant to current or future litigation.[23] Once litigation is reasonably anticipated, a company should establish and follow a reasonable preservation plan.[24]\u00a0 Although there are no specific court-sanctioned processes for complying with the preservation duty, courts generally measure the parties\u2019 conduct in a given case against the standards of reasonableness and good faith.[25]\u00a0 In this context, a \u201cdefined policy and memorialized evidence of compliance should provide strong support if the organization is called up on to prove the reasonableness of the decision-making process.\u201d[26]<\/p>\n<p>[17]\u00a0\u00a0\u00a0\u00a0\u00a0 The duty to preserve is not without limits: \u201c[e]lectronic discovery burdens should be proportional to the amount in controversy and the nature of the case\u201d so the high cost of electronic discovery does not \u201coverwhelm the ability to resolve disputes fairly in litigation.\u201d[27]\u00a0Moreover, courts do not equate reasonableness with \u201cperfection.\u201d[28] Nor does the law require a party to take \u201cextraordinary\u201d measures to preserve \u201cevery e-mail\u201d even if it is technically feasible to do so.[29]\u00a0 \u201cRather, in accordance with existing records and information management principles, it is more rational to establish a procedure by which selected items of value can be identified and maintained as necessary to meet the organization\u2019s legal and business needs[.]\u201d[30]<\/p>\n<p>[18]\u00a0\u00a0\u00a0\u00a0\u00a0 Critical tasks in a preservation plan are the identification and documentation of key custodians and other sources of potentially relevant information.[31] Custodians identified as having potentially relevant information should generally receive a written litigation hold notice.[32]\u00a0 The notice should be sent by someone occupying a position of authority within the organization to increase the likelihood of compliance.[33] The Sedona Guidelines also suggest that a hold notice is most effective when it:<\/p>\n<p style=\"padding-left: 30px\"><em>\u00a01)\u00a0\u00a0\u00a0\u00a0\u00a0 Identifies the persons likely to have relevant information and communicates a preservation notice to those persons;<\/em><\/p>\n<p style=\"padding-left: 30px\"><em>2)\u00a0\u00a0\u00a0\u00a0\u00a0 Communicates the preservation notice in a manner that ensures the recipients will receive actual, comprehensible and effective notice of the requirement to preserve information;<\/em><\/p>\n<p style=\"padding-left: 30px\"><em>3)\u00a0\u00a0\u00a0\u00a0\u00a0 Is in written form;<\/em><\/p>\n<p style=\"padding-left: 30px\"><em>4)\u00a0\u00a0\u00a0\u00a0\u00a0 Clearly defines what information is to be preserved and how the preservation is to be undertaken; and<\/em><\/p>\n<p style=\"padding-left: 30px\"><em>5)\u00a0\u00a0\u00a0\u00a0\u00a0 Is regularly reviewed and reissued in either its original form or an amended form when necessary.[34]<\/em><\/p>\n<p>\u00a0[19]\u00a0\u00a0\u00a0\u00a0\u00a0 The legal hold should also include a mechanism for confirming that recipients received and understood the notice, for following up with custodians who do not acknowledge receipt, and for escalating the issue until it is resolved.[35]\u00a0 To be effective, the legal hold should be periodically reissued to remind custodians of their obligation and to apprise them of changes required by the facts and circumstances in the litigation.[36]<\/p>\n<p>[20]\u00a0\u00a0\u00a0\u00a0\u00a0 Experience has also shown that legal holds that are not properly managed and ultimately released are less likely to receive the appropriate level of attention by employees. Thus, the legal hold process should also include a means for determining when litigation is no longer reasonably anticipated and the hold can be released, while ensuring that information relevant to another active matter is preserved.[37]<\/p>\n<h3>\u00a0<b>B.\u00a0 The Remediation Framework<\/b><\/h3>\n<p><b>\u00a0<\/b>[21]\u00a0\u00a0\u00a0\u00a0\u00a0 Against this backdrop, it is possible to outline a framework for data remediation that is compliant with legal preservation requirements.\u00a0 The following describes a high-level data remediation process that can be applied to virtually any data environment and any risk tolerance profile.\u00a0 The general process is described in Figure 1 below:<\/p>\n<p><a href=\"https:\/\/jolt.richmond.edu\/files\/2014\/03\/Figure-1-Kiker-e1394813470271.jpg\"><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-2023 aligncenter\" alt=\"Figure 1 - Kiker\" src=\"https:\/\/jolt.richmond.edu\/files\/2014\/03\/Figure-1-Kiker-e1394813470271.jpg\" width=\"1000\" height=\"659\" \/><\/a><\/p>\n<p align=\"center\">Figure 1: Data Remediation Framework<\/p>\n<h4 style=\"padding-left: 30px\"><b>1.\u00a0 Assemble the Team<\/b><\/h4>\n<p><b>\u00a0<\/b>[22]<b>\u00a0\u00a0\u00a0\u00a0\u00a0 <\/b>A successful data remediation project depends on invested participation by at least three constituents in the organization: legal, information technology (\u201cIT\u201d), and records and information management (\u201cRIM\u201d).\u00a0 In addition, the project may require additional support from experts experienced in information search and retrieval and statistical analysis.\u00a0 In-house and\/or outside counsel provides legal oversight and risk assessment for the project team, as well as guidance on legal preservation obligations.\u00a0 IT provides the technological expertise necessary to understand the structure and capabilities of the target data repository.\u00a0 RIM professionals provide guidance on business and regulatory retention obligations.\u00a0 The need for information search and retrieval experts and statisticians depends on the complexity of the data remediation effort as described below.\u00a0 Finally, including business users of the information may be necessary as required to fully document retention requirements applicable to a particular repository if not adequately documented in the organization\u2019s document retention policy and schedule.<\/p>\n<h4 style=\"padding-left: 30px\">\u00a0<b>2.\u00a0 Select Target Data Repository<\/b><\/h4>\n<p><b>\u00a0<\/b>[23]\u00a0\u00a0\u00a0\u00a0\u00a0 Selecting the target data repository requires consideration of the costs and benefits of the data remediation exercise.\u00a0 Each type of repository presents unique opportunities and challenges.\u00a0 For example, e-mail systems, whether traditional or archived, are notorious for containing vast amounts of information that is not needed for any business or legal purpose.\u00a0 Similarly, shared network drives tend to contain large volumes of unused and unneeded information.\u00a0 Backup tapes, legacy systems, and even structured databases are other possible targets.\u00a0 IT and RIM resources are invaluable in identifying a suitable target repository.\u00a0 For example, IT can often run reports identifying directories and files that have not been accessed recently.<\/p>\n<h4 style=\"padding-left: 30px\">\u00a0<b>3.\u00a0 Document Retention and Preservation Obligations<\/b><\/h4>\n<p>[24]\u00a0\u00a0\u00a0\u00a0\u00a0 As discussed above, it is critical to understand the retention and preservation obligations that are applicable to the data contained in the target repository.\u00a0 Retention obligations include the business information needs as well as any regulatory requirements mandating the preservation of data.\u00a0 Ideally, these are incorporated into the document retention policy and schedule for the organization.\u00a0 If not, it will be important to document those requirements applicable to the target repository.<\/p>\n<p>[25]\u00a0\u00a0\u00a0\u00a0\u00a0 Preservation obligations are driven by existing and reasonably anticipated litigation.[38]\u00a0 In some cases this may be the most challenging part of the project, particularly for highly litigious companies, because, unlike business needs and regulatory requirements, preservation obligations are constantly changing as new matters arise and circumstances evolve in existing matters.\u00a0 Successful completion of the remediation project will require a detailed understanding of, and constant attention to, the preservation obligations applicable to the target repository.\u00a0 As discussed below, some of the risk associated with this aspect of the project can be ameliorated through selection of the appropriate repository and culling criteria.\u00a0 Nevertheless, the scope and timing of the project will be driven in large part by the preservation obligations applicable to the target repository.<\/p>\n<h4 style=\"padding-left: 30px\"><b>4.\u00a0 Inventory Target Data Repository<\/b><\/h4>\n<p>[26]\u00a0\u00a0\u00a0\u00a0\u00a0 After selecting the target data repository, the team must inventory the information within that repository.\u00a0 This does not involve creating an exhaustive list or catalog of every item within the repository.\u00a0 Rather, inventorying the repository involves developing a good understanding of the types of information that are contained there, the date ranges of the information, and other criteria that will enable identifying information that must be retained and that which can be deleted.\u00a0 The details of the inventory will vary by data repository.\u00a0 For example, for an e-mail server, the pertinent criteria may include only date ranges and custodians, whereas for a shared network drive, the pertinent criteria may include departments and individuals with access, date ranges, and file types.<\/p>\n<h4 style=\"padding-left: 30px\">\u00a0<b>5.\u00a0 Gross Culling<\/b><\/h4>\n<p><b>\u00a0<\/b>[27]\u00a0\u00a0\u00a0\u00a0\u00a0 The next step is to determine the \u201cgross culling\u201d criteria for the data repository.\u00a0 In this context, \u201cgross culling\u201d refers to an initial phase of data culling based on broad criteria as opposed to fine or detailed culling criteria that may be used in a later phase of the exercise.[39]\u00a0 The nature of the information contained within the repository will determine the specific criteria to be used, but the objective is to locate the \u201clow-hanging fruit,\u201d the items within the repository that can be readily identified as not falling within any retention or preservation obligation. These are black-and-white decisions where the remediation team can definitively determine without further analysis that the items identified can be deleted.<\/p>\n<p>[28]\u00a0\u00a0\u00a0\u00a0\u00a0 For example, in most cases, dates are effective gross culling criteria.\u00a0 Quite often, large volumes of e-mail and loose files (data retained in shared network drives or other unstructured storage) predate any existing retention or preservation obligation for such items.\u00a0 Similarly, in repositories that are subject to short or no retention guidelines, the business need for the data can be evaluated in terms of the date last accessed.\u00a0 In the case of shared network drives, for example, it is not uncommon to find large volumes of information that has not been accessed by any user in many years.[40]\u00a0 Such information can be disposed of with very little risk.<\/p>\n<h4 style=\"padding-left: 30px\">\u00a0<b>6.\u00a0 Fine Culling<\/b><\/h4>\n<p><b>\u00a0<\/b>[29]\u00a0\u00a0\u00a0\u00a0\u00a0 Sometimes, the process need go no further than the gross culling stage.\u00a0 Depending on the volume of data deleted and the volume and nature of the data remaining, the remediation team may determine that the cost and benefit of attempting further culling of the data are not worth the effort and risk.\u00a0 In some cases, however, gross culling techniques will not identify sufficient volumes of unneeded data and more sophisticated culling strategies must be employed.<\/p>\n<p>[30]\u00a0\u00a0\u00a0\u00a0\u00a0 The precise culling technique and strategy will depend on the specific data repository, its native search capabilities, and the availability of other search tools.\u00a0 For example, many modern e-mail archiving systems have fairly sophisticated native search capabilities that can locate with a high degree of accuracy content pertinent to selected criteria.\u00a0 Other systems will require the use of third-party technology.\u00a0 In either case, the fine culling process will require selection of culling criteria that will uniquely identify items not subject to a retention or preservation obligation and be susceptible to verification.\u00a0 Depending on the nature of the data and the complexity of the necessary search criteria, the remediation team may need to engage an expert in information search and retrieval.<\/p>\n<h4 style=\"padding-left: 30px\">\u00a0<b>7.\u00a0 Sampling and Statistical Analysis<\/b><\/h4>\n<p>[31]\u00a0\u00a0\u00a0\u00a0\u00a0 Regardless of the specific fine culling strategy employed, the remediation team should validate the results by sampling and analysis to ensure defensibility.\u00a0 Generally, it will be advisable to engage a statistician to direct the sampling effort and perform the analysis because both can be quite complex and rife with opportunity for error.[41]\u00a0 Moreover, in the event that the company\u2019s process is ever challenged, validation by an independent expert is compelling evidence of good faith.\u00a0 It is important to realize that the statistical analysis cannot demonstrate that no items subject to a preservation obligation are included in the data to be destroyed.\u00a0 It can only identify the probability that this is the case, but it can do so with remarkable precision when properly performed.[42]<\/p>\n<h4 style=\"padding-left: 30px\">\u00a0<b>8.\u00a0 Iteration<\/b><\/h4>\n<p>[32]\u00a0\u00a0\u00a0\u00a0\u00a0 Fine culling and validation should continue until the remediation team achieves results that meet its expectations regarding the volume of data identified for deletion and the probability that only data not subject to a preservation obligation are included in the result set.<\/p>\n<p>&nbsp;<\/p>\n<h2 style=\"text-align: center\"><b>\u00a0<\/b><b>IV.\u00a0 Conclusion<\/b><\/h2>\n<p style=\"text-align: left\" align=\"center\">[33]\u00a0\u00a0\u00a0\u00a0\u00a0 The enormity of the challenge that expanding volumes of unneeded information creates for businesses is difficult to understate.\u00a0 Companies literally spend millions of dollars annually to store and maintain information that serves no useful purpose, funds that could be directed to productive uses such as hiring, research, and investment.\u00a0 Facing this challenge, on the other hand, is a challenge of its own, perhaps due more to the fear of adverse consequences in litigation than any other factor.\u00a0 It is possible, however, to develop a defensible data remediation process that enables a company to demonstrate good faith and reasonableness while eliminating the cost, waste, and risk of this unnecessary data.<\/p>\n<p>&nbsp;<\/p>\n<div><\/p>\n<hr align=\"left\" size=\"1\" width=\"33%\" \/>\n<div>\n<p>* Dennis Kiker has been a partner in a national law firm, director of professional services at a major e-Discovery company, and a founding shareholder of his own law firm. He has served as national discovery counsel for one of the largest manufacturing companies in the country, and counseled many others on discovery and information governance-related issues. He is a Martindale-Hubbell AV-rated attorney admitted at various times to practice in Virginia, Arizona and Florida, and holds a J.D., <i>magna cum laude<\/i> &amp; Order of the Coif from the University of Michigan Law School.\u00a0 Dennis is currently a consultant at Granite Legal Systems, Inc. in Houston, Texas.<\/p>\n<p>&nbsp;<\/p>\n<\/div>\n<div>\n<p>[1] <i>See <\/i>The Sedona Conference, The Sedona Principles: Second Edition Best Practices Recommendations &amp; Principles for Addressing Electronic Document Production\u00a0 28 (Jonathan M. Redgrave et al. eds., 2007) [hereinafter \u201cThe Sedona Principles\u201d], <i>available at<\/i> http:\/\/www.sos.mt.gov\/Records\/committees\/erim_resources\/A%20-%20Sedona%20Principles%20Second%20Edition.pdf (last visited Jan. 30, 2014);<i> see also<\/i> Louis R. Pepe &amp; Jared Cohane<i>, Document Retention, Electronic Discovery, E-Discovery Cost Allocation, and Spoliation Evidence: The Four Horsemen of the Apocalypse of Litigation Today<\/i>, 80 Conn. B. J. 331, 348 (2006) (explaining how proposed Rule 37(f) addresses the routine alteration and deletion of electronically stored information during ordinary use).<\/p>\n<p>[2] <i>See <\/i>The Sedona Principles<i>, supra <\/i>note 1, at 12.<\/p>\n<div>\n<p>[3] <i>See <\/i>Peter Lyman &amp; Hal R. Varian, <i>How Much Information 2003?<\/i>,<i> <\/i>http:\/\/www.sims.berkeley.edu\/research\/projects\/how-much-info-2003\/ (last visited Feb. 9, 2014).<\/p>\n<\/div>\n<div>\n<p>[4] <i>Id.<\/i><\/p>\n<\/div>\n<div>\n<p>[5] <i>See id.<\/i><\/p>\n<\/div>\n<div>\n<p>[6] Jake Frazier, <i>Hoarders: The Corporate Edition<\/i>, Business Computing World\u00a0 (Sept. 25, 2013), http:\/\/www.businesscomputingworld.co.uk\/hoarders-the-corporate-edition\/.<\/p>\n<\/div>\n<div>\n<p>[7] <i>Id.<\/i><\/p>\n<\/div>\n<div>\n<p><i> <\/i>[8] <i>See<\/i> James Dertouzos et. al, Rand Inst. for Civil Justice, The Legal and Economic Implications of E-Discovery: Options for Future Research ix (2008), <i>available at<\/i> http:\/\/www.rand.org\/content\/dam\/rand\/pubs\/occasional_papers\/2008\/RAND_OP183.pdf; <i>see also <\/i>Robert Blumberg &amp; Shaku Atre, <i>The Problem with Unstructured Data<\/i>, Info. Mgmt. (Feb. 1, 2003, 1:00 AM), http:\/\/soquelgroup.com\/Articles\/dmreview_0203_problem.pdf; The Radicati Group, Taming the Growth of Email: An ROI Analysis 3-4 (2005), <i>available at <\/i>http:\/\/www.radicati.com\/wp\/wp-content\/uploads\/2008\/09\/hp_whitepaper.pdf<\/p>\n<\/div>\n<div>\n<p>[9] <i>See <\/i>David C. Blair &amp; M.E. Maron, <i>An Evaluation of Retrieval Effectiveness for a Full-Text Document Retrieval <\/i>System, Comm. ACM, March 1985, at 289-90, 295-96.<\/p>\n<p>[10] <i>See <\/i>Ellen M. Voorhees, <i>Variations in Relevance Judgments and the Measurement of Retrieval Effectiveness<\/i>, 36 Info. Processing &amp; Mgmt. 697, 701 (2000), <i>available at<\/i> http:\/\/\u200cwww.cs.cornell.edu\/\u200ccourses\/\u200ccs430\/\u200c2006fa\/\u200ccache\/\u200cTrec_8.pdf (finding that relevance is not a consistently applied concept between independent reviewers).\u00a0 <i>See generally <\/i>Hebert L. Roitblat et al., <i>Document Categorization in Legal Electronic Discovery: Computer Classification vs. Manual Review<\/i>, 61 J. Am. Soc\u2019y. for Info. Sci. &amp; Tech. 70, 77 (2010).<\/p>\n<\/div>\n<div>\n<p>[11] <i>See <\/i>Voorhees, <i>supra<\/i> note 10, at 701 (finding that the \u201coverlap\u201d between even senior reviewers shows that they disagree as often as they agree on relevance).<\/p>\n<\/div>\n<div>\n<p>[12]\u00a0 <i>See generally <\/i>Maura R. Grossman &amp; Gordon V. Cormack, <i>Technology-Assisted Review in E-Discovery Can Be More Effective and More Efficient Than Exhaustive Manual Review<\/i>, 17 Rich. J.L. &amp; Tech. 11 \u00b6 2 (2011),<i> <\/i>http:\/\/\u200cjolt.\u200crichmond.\u200cedu\/\u200c\u200cv17i3\/\u200carticle11.pdf (analyzing data from the TREC 2009 Legal Track Interactive Task Initiative).<\/p>\n<\/div>\n<div>\n<p>[13] <i>See <\/i>Moore v. Publicis Groupe SA, No. 11 Civ. 1279(ALC)(AJP), 2012 WL 1446534, at *1-3 (S.D.N.Y. Apr. 26, 2012).<\/p>\n<\/div>\n<div>\n<p>[14] <i>See <\/i>Global Aerospace, Inc. v. Landow Aviation, L.P<i>.<\/i>, No. CL 61040, 2012 Va. Cir. LEXIS 50, at *2 (Va. Cir. Ct. Apr. 23, 2012).<\/p>\n<\/div>\n<div>\n<p>[15] <i>See <\/i>Mem. in Supp. of Mot. for Protective Order Approving the Use of Predictive Coding at 4-5, Global Aerospace, Inc. v. Landow Aviation, L.P., No. CL 61040, 2012 Va. Cir. LEXIS 50 (Va. Cir. Ct. Apr. 9, 2012).<\/p>\n<\/div>\n<div>\n<p>[16] <i>Id.<\/i> at 6-7.<\/p>\n<\/div>\n<div>\n<p>[17] Fed. R. Evid. 502(b) Advisory Committee\u2019s Notes, Subdivision (b) (2007).<\/p>\n<\/div>\n<div>\n<p>[18] Arthur Anderson LLP v. United States, 544 U.S. 696, 704 (2005).<\/p>\n<\/div>\n<div>\n<p>[19] <i>Id.<\/i>; <i>see<\/i> Managed Care Solutions, Inc. v. Essent Healthcare, 736 F. Supp. 2d 1317, 1326 (S.D. Fla. 2010) (rejecting plaintiffs\u2019 argument that a company policy that e-mail data be deleted after 13 months was unreasonable) (citing Wilson v. Wal-Mart Stores, Inc<i>.<\/i>, No. 5:07-cv-394-Oc-10GRJ, 2008 WL 4642596, at *2 (M.D. Fla. Oct. 17, 2008); Floeter v. City of Orlando, No. 6:05-CV-400-Orl-22KRS, 2007 WL 486633, at *7 (M.D. Fla. Feb. 9, 2007)).\u00a0 <i>But see <\/i>Day v. LSI Corp<i>., <\/i>No. CIV 11\u2013186\u2013TUC\u2013CKJ, 2012 WL 6674434, at *16 (D. Ariz. Dec. 20, 2012) (finding evidence of defendant\u2019s failure to follow its own document policy was a factor in entering default judgment sanction for spoliation).<\/p>\n<\/div>\n<div>\n<p><i> <\/i>[20] For purposes of this article, such laws and regulations are treated as retention requirements with which a business must comply in the ordinary course of business.\u00a0 This article focuses on the requirement to exempt records from \u201cordinary course\u201d retention requirements due to a duty to preserve the records when litigation is reasonably anticipated.\u00a0 In short, this article relies on the distinction between <i>retention<\/i> of information and <i>preservation <\/i>of information, focusing on the latter.\u00a0 <i>See<\/i><i>infra <\/i>text accompanying note 23.<\/p>\n<\/div>\n<div>\n<p>[21] <i>See <\/i>Sylvestri v. Gen. Motors, Inc<i>.<\/i>, 271 F.3d 583, 590 (4th Cir. 2001); <i>see also<\/i> Stanley, Inc. v. Creative Pipe, Inc.<i>,<\/i> 269 F.R.D. 497, 519 (4th Cir. 2010).<\/p>\n<\/div>\n<div>\n<p>[22] <i>See <\/i>Cache la Poudre Feeds v. Land O\u2019Lakes, 244 F.R.D. 614, 621, 623 (D. Colo. 2007); <i>see also<\/i> The Sedona Principles<i>, supra <\/i>note 1, at 14.<\/p>\n<\/div>\n<div>\n<p>[23] <i>See <\/i><em>Pension Comm. of the Univ. of Montreal Pension Plan v. Banc of Am. Sec., LLC, <\/em>685 F. Supp. 2d 456, 466 (S.D.N.Y. Jan. 15, 2010 <i>as amended<\/i> May 28, 2010); Rimkus Consulting Grp., Inc. v. Cammarata, 688 F. Supp. 2d 598, 612-13 (S.D. Tex. 2010);\u00a0 Zubulake v. UBS Warburg LLC, 220 F.R.D. 212, 216 (S.D.N.Y. 2003) (<i>Zubulake IV<\/i>); <i>see also<\/i> The Sedona Conference, <i>Commentary on Legal Holds: The Trigger &amp; The Process<\/i>, 11 Sedona Conf. J. 265, 269 (2010) [hereinafter \u201c<i>Commentary on Legal Holds<\/i>\u201d].<\/p>\n<\/div>\n<div>\n<p>[24] <i>Commentary on Legal Holds<\/i>, <i>supra <\/i>note 23,<i> <\/i>at 269 (\u201cAdopting and consistently following a policy or practice governing an organization\u2019s preservation obligations are factors that may demonstrate reasonableness and good faith.\u201d); <i>see <\/i>The Sedona Principles, <i>supra<\/i> note 1, at 12.<\/p>\n<\/div>\n<div>\n<p>[25] <i>Commentary on Legal Holds<\/i>, <i>supra<\/i> note 23, at 270 (evaluating an organization\u2019s preservation decisions should be based on good faith and reasonable evaluation of relevant facts and circumstances).<\/p>\n<\/div>\n<div>\n<p>[26] <i>Id.<\/i> at 274.<\/p>\n<p>[27] <em>Rimkus Consulting<\/em>, 688 F. Supp. 2d at 613 n.8 (quoting The Sedona Principles, <i>supra <\/i>note 1, at 17); <i>see also <\/i>Stanley v. Creative Pipe, Inc<i>.<\/i>, 269 F.R.D. 497, 523 (D. Md., 2010); <i>Commentary on Legal Holds<\/i>, <i>supra<\/i> note 23, at 270.<\/p>\n<\/div>\n<div>\n<p>[28] <em>Pension Comm.<\/em><em>,<\/em> 685 F. Supp. 2d at 461 (\u201cCourts cannot and do not expect that any party can meet a standard of perfection.\u201d).<\/p>\n<\/div>\n<div>\n<p>[29] <i>See <\/i>The Sedona Principles, <i>supra<\/i> note 1, at 28, 30 (citing Concord Boat Corp. v. Brunswick Corp<i>.<\/i>, No. LR-C-95-781, 1997 WL 33352759, at *4 (E.D. Ark. Aug. 29, 1997)).<\/p>\n<\/div>\n<div>\n<p>[30] The Sedona Principles, <i>supra<\/i> note 1, at 15.<\/p>\n<\/div>\n<div>\n<p>[31] <i>See Commentary on Legal Holds<\/i>,<i> supra<\/i> note 23, at 270;<i> id.<\/i> at 28.<\/p>\n<\/div>\n<div>\n<p>[32] <i>See <em>Pension Comm. <\/em><\/i>685 F. Supp. 2d at 465; s<i>ee also Commentary on Legal Holds<\/i>, <i>supra<\/i> note 23, at 270.<\/p>\n<\/div>\n<div>\n<p>[33] The Sedona Principles, <i>supra<\/i>, note 1, at 32.<\/p>\n<\/div>\n<div>\n<p>[34] <i>Commentary on Legal Holds<\/i>, <i>supra<\/i> note 23, at 270.<\/p>\n<\/div>\n<div>\n<p>[35] <i>Id.<\/i> at 283-85.<\/p>\n<\/div>\n<div>\n<p>[36] <i>S<\/i><i>ee id. <\/i>at 285.<\/p>\n<\/div>\n<div>\n<p>[37] <i>Id. <\/i>at 287.<\/p>\n<\/div>\n<div>\n<p>[38] <i>See supra<\/i> \u00b6 16.<\/p>\n<\/div>\n<div>\n<p>[39] <i>See <\/i>Alex Vorro, <i>How to Reduce Worthless Data<\/i>,<i> <\/i>InsideCounsel (Mar. 1, 2012), http:\/\/www.insidecounsel.com\/2012\/03\/01\/how-to-reduce-worthless-data?t=technology.<\/p>\n<p>[40] <i>See, e.g.<\/i>, Anne Kershaw, <i>Hoarding Data Wastes Money<\/i>, Baseline (Apr. 16, 2012), http:\/\/www.baselinemag.com\/storage\/Hoarding-Data-Wastes-Money\/ (80% of the data on shared network and local hard drives has not been accessed in three to five years).<\/p>\n<\/div>\n<div>\n<p>[41] Statistical sampling results can be as valid using a small random sample size as they are for using a larger sample size because, in a simple random sample of any given size, all items are given an equal probability of being selected for the statistical assessment.\u00a0 In fact, to achieve a confidence interval of 95% with a margin of error of 5%, a sample size of 384 would be sufficient for the population of 300 million.\u00a0 <i>See<\/i><i>Sample Size Table<\/i>, Research Advisors, http:\/\/research-advisors.com\/tools\/SampleSize.htm (last visited on Jan. 12, 2014) (citing Robert V. Krejcie &amp; Daryle W. Morgan, <i>Determining Sample Size for Research Activities<\/i>, <em>Educational and Psychological Measurement<\/em> 30 Educ. &amp; Psychol. Measurement 607, 607-610 (1970).\u00a0 However, samples can be vulnerable to discrete \u201csampling error\u201d because the randomness of the selection may result in a sample that does not reflect the makeup of the overall population.\u00a0 For instance, a simple random sample of messages will on average produce five with attachments and five with no attachments, but any given test may over-represent one message type (e.g., those with attachments) and under-represent the other (e.g., those without).<\/p>\n<\/div>\n<div>\n<p>[42] <i>See,<\/i><i>e.g.<\/i>, <i>Statistics<\/i>, Wikipedia, http:\/\/en.wikipedia.org\/wiki\/Statistics (last visited on Feb. 9, 2014).<\/p>\n<\/div>\n<div><\/div>\n<\/div>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>DownloadPDF Cite as: Dennis R. Kiker,\u00a0Defensible Data Deletion: A Practical Approach to Reducing Cost and Managing Risk Associated with Expanding Enterprise Data, 20 Rich. J.L. &amp; Tech. 6 (2014), http:\/\/jolt.richmond.edu\/v20i2\/article6.pdf. \u00a0 Dennis R. Kiker* \u00a0 I.\u00a0 Introduction [1]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 Modern businesses are hosts to steadily increasing volumes of data, creating significant cost and risk while potentially [&hellip;]<\/p>\n","protected":false},"author":4287,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"jetpack_post_was_ever_published":false,"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":false,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2}},"categories":[1228],"tags":[],"class_list":["post-2021","post","type-post","status-publish","format-standard","hentry","category-articles"],"jetpack_publicize_connections":[],"jetpack_featured_media_url":"","jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/paMHOZ-wB","jetpack-related-posts":[],"_links":{"self":[{"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/posts\/2021","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/users\/4287"}],"replies":[{"embeddable":true,"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/comments?post=2021"}],"version-history":[{"count":0,"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/posts\/2021\/revisions"}],"wp:attachment":[{"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/media?parent=2021"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/categories?post=2021"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/tags?post=2021"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}