{"id":2017,"date":"2014-03-14T16:41:16","date_gmt":"2014-03-14T16:41:16","guid":{"rendered":"http:\/\/jolt.richmond.edu\/?p=2017"},"modified":"2019-03-08T19:52:30","modified_gmt":"2019-03-09T00:52:30","slug":"finding-the-signal-in-the-noise-information-governance-analytics-and-the-future-of-legal-practice","status":"publish","type":"post","link":"https:\/\/blog.richmond.edu\/jolt\/2014\/03\/14\/finding-the-signal-in-the-noise-information-governance-analytics-and-the-future-of-legal-practice\/","title":{"rendered":"Finding the Signal in the Noise: Information Governance, Analytics, and the Future of Legal Practice"},"content":{"rendered":"<p style=\"text-align: left\" align=\"center\"><a href=\"http:\/\/jolt.richmond.edu\/v20i2\/article7.pdf\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-128\" alt=\"pdf_icon\" src=\"http:\/\/jolt.richmond.edu\/files\/2012\/05\/pdf_icon1.gif\" width=\"16\" height=\"16\" \/>DownloadPDF<\/a><\/p>\n<p align=\"center\">Cite as: Bennett B. Borden &amp; Jason R. Baron,\u00a0<i>Finding the Signal in the Noise: Information Governance, Analytics, and the Future of Legal Practice<\/i>, 20 Rich. J.L. &amp; Tech. 7 (2014), http:\/\/jolt.richmond.edu\/v20i2\/article7.pdf.<\/p>\n<p align=\"center\"><b>\u00a0<\/b><\/p>\n<p align=\"center\">Bennett B. Borden* and Jason R. Baron**<\/p>\n<p align=\"center\"><b>\u00a0<\/b><\/p>\n<h2 align=\"center\"><b>Introduction<\/b><\/h2>\n<p>[1]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 In the watershed year of 2012, the world of law witnessed the first concrete discussion of how predictive analytics may be used to make legal practice more efficient.\u00a0 That the conversation about the use of predictive analytics has emerged out of the e-Discovery sector of the law is not all that surprising: in the last decade and with increasing force since 2006\u2014with the passage of revised Federal Rules of Civil Procedure that expressly took into account the fact that lawyers must confront \u201celectronically stored information\u201d in all its varieties\u2014there has been a growing recognition among courts and commentators that the practice of litigation is changing dramatically.\u00a0 What needs now to be recognized, however, is that the rapidly evolving tools and techniques that have been so helpful in providing efficient responses to document requests in complex litigation may be used in a variety of complementary ways to the discovery process itself.<\/p>\n<p>[2]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 This Article is informed by the authors\u2019 strong views on the subject of using advanced technological strategies to be better at \u201cinformation governance,\u201d as defined herein.\u00a0 If a certain evangelical strain appears to arise out of these pages, the authors willingly plead guilty.\u00a0 One need not be an evangelist, however, but merely a realist to recognize that the legal world and the corporate world both are increasingly confronting the challenges and opportunities posed by \u201cBig data.\u201d[1]\u00a0 This Article has a modest aim: to suggest certain paths forward where lawyers may add value in recommending to their clients greater use of advanced analytical techniques for the purpose of optimizing various aspects of information governance.\u00a0 No attempt at comprehensiveness is aimed for here; instead, the motivation behind writing this Article is simply to take stock of where the legal profession is, as represented by the emerging case law on predictive coding represented by <i>Da Silva Moore<\/i>,[2] and to suggest that the expertise law firms have gained in this area may be applied in a variety of related contexts.<\/p>\n<p>[3]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 To accomplish what we are setting out to do, we will divide the discussion into the following parts: first, a synopsis of why and how predictive coding first emerged against the backdrop of e-Discovery.\u00a0 This discussion will include a brief overview of predictive coding with references to the technical literature, as the subject has been recently covered exhaustively elsewhere.\u00a0 Second, we will define what we mean by \u201cBig data,\u201d \u201canalytics,\u201d and \u201cinformation governance,\u201d for the purpose of providing a proper context for what follows.\u00a0 Third, we will note those aspects of an information governance program that are most susceptible to the application of predictive coding and related analytical techniques.\u00a0 Perhaps of most value, we wish to share a few \u201cearly\u201d examples of where we as lawyers have brought advanced analytics, like predictive coding, to bear in non-litigation contexts and to assist our clients in creative new ways.\u00a0 We fully expect that what we say here will be overrun with a multitude of real-life use cases soon to emerge in the legal space. \u00a0Armed with the knowledge that we are attempting to catch lightning in a bottle and that law reviews on subjects such as this one have ever decreasing \u201cshelf-lives\u201d[3] in terms of the value proposition they provide, we proceed nonetheless.<\/p>\n<p><b>A.\u00a0 The Path to <i>Da Silva Moore<\/i><\/b><\/p>\n<p>[4]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 <em>The Law of Search and Retrieval.<\/em>\u00a0 In the beginning, there was manual review.\u00a0\u00a0Any graduate of a law school during the latter part of the twentieth century who found herself or himself employed before the year 2000 at a law firm specializing in litigation and engaged in high-stakes discovery remembers well how document review was conducted: legions of lawyers with hundreds if not thousands of boxes in warehouses, reviewing folders and pages one-by-one in an effort to find the relevant needles in the haystack.[4]\u00a0 (Some of us also remember \u201cSheparding\u201d a case to find subsequent citations to it, using red and yellow booklets, before automated key-citing came along.)\u00a0 Although manual review continues to remain a default practice in a variety of more modest engagements, it is increasingly the case that all of discovery involves \u201ce-Discovery\u201d of some sort\u2014that the world is simply \u201cawash in data\u201d[5] (starting but by no means ending with email, messages and other textual documents of all varieties), and that it will increasingly be the unusual case of any size where documents in paper form still loom large as the principal source of discovery.<\/p>\n<p>[5]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 At the turn of the century, the dawning awareness of the need to deal with a new realm of electronically stored information (\u201cESI\u201d) led to burgeoning efforts on many fronts, including, for example, the creation of The Sedona Conference working group on electronic document retention and production, members of which drafted <i>The Sedona Conference Principles: Addressing Electronic Document Production <\/i>(2005; 2d ed. 2007) and its \u201cprequel,\u201d <i>The Sedona Guidelines: Best Practice Guidelines and Commentary for Managing Records and Information in the Electronic Age <\/i>(2005; 2d ed. 2007).\u00a0 These early commentaries, including a smattering of pre-2006 case law,[6] recognized that changes in legal practice were necessary to accommodate the big changes coming in the world of records and information management within the enterprise.\u00a0 Subsequent developments would constitute various complementary threads leading to the greater use of analytics in the legal space.<\/p>\n<p>[6]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 First, part of that early recognition was that in an inflationary universe of rapidly expanding amounts of ESI, new tools and techniques would be necessary for the legal profession to adapt and keep up with the times.[7]\u00a0 By the time of adoption of the revised Federal Rules of Civil Procedure in 2006, which expressly added the term \u201cESI\u201d to supplement \u201cdocuments\u201d in the rule set applicable to discovery practice, the legal profession was well aware of the need to perform automated searches in the form of keyword searching within large data sets as the only realistically available means for sorting information into relevant and non-relevant evidence in particular engagements, be they litigation or investigations.\u00a0 So too, it was recognized early on in commentaries[8] and followed by case law[9] that keyword searching, as good a tool as it was, had profound limitations that in the end do not scale well.\u00a0 At the end of the day, even being able to limit or cull down a large data set to one percent of its original size through the use of keywords leaves the lawyer with the near impossible task of manually reviewing a very large set of documents at great cost.[10]<\/p>\n<p>[7]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 Second, in evolving e-Discovery practice after 2006, a growing recognition also occurred around the idea that e-Discovery workflows are an \u201cindustrial\u201d process in need of better metrics and measures for evaluating the quality of productions of large data sets.\u00a0 As recognized in <i>The Sedona Conference Commentary on Achieving Quality in E-discovery <\/i>(Post-Public Comment Version 2013):<\/p>\n<p style=\"padding-left: 30px\"><em>The legal profession has passe<br \/>\nd a crossroads: When faced with a choice between continuing to conduct discovery as it had \u201calways been practiced\u201d in a paper world\u2014before the advent of computers, the Internet, and the exponential growth of electronically stored information (ESI)\u2014or alternatively embracing new ways of thinking in today\u2019s digital world, practitioners and parties acknowledged a new reality and chose progress.\u00a0 But while the initial steps are completed, cost-conscious clients and over-burdened judges are increasingly demanding that parties find new approaches to solve litigation problems.[11]<\/em><\/p>\n<p>\u00a0[8]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 The Commentary goes on to suggest that the legal profession would benefit from greater<\/p>\n<p style=\"padding-left: 30px\"><em>awareness about a variety of processes, tools, techniques, methods, and metrics that fall broadly under the umbrella term \u201cquality measures\u201d and that may be of assistance in handling ESI throughout the various phases of the discovery workflow process. \u00a0These include greater use of project management, sampling, machine learning, and other means to verify the accuracy and completeness of what constitutes the \u201coutput\u201d of e-[D]iscovery. \u00a0Such collective measures, drawn from a wide variety of scientific and management disciplines, are intended only as an entry-point for further discussion, rather than an all-inclusive checklist or cookie-cutter solution to all e-[D]iscovery issues.[12]<\/em><\/p>\n<p>\u00a0[9]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 Indeed, more recent case law has recognized the need for quality control, including through the use of greater sampling, iterative methods, and phased productions in line with principles of proportionality.[13]\u00a0 Still other case law has emphasized the need for cooperation among parties in litigation on technical subjects, especially at the margins of, or outside the range of, lawyer expertise if not basic competence.<\/p>\n<p>[10]\u00a0\u00a0\u00a0\u00a0\u00a0 Active or supervised \u201cmachine learning,\u201d as referred to here in the context of e-Discovery, refers to a set of analytical tools and techniques that go by a variety of names, such as \u201cpredictive coding,\u201d \u201ccomputer-assisted review,\u201d and \u201ctechnology assisted review.\u201d\u00a0 As explained in one helpful recent monograph:<\/p>\n<p style=\"padding-left: 30px\"><em>Predictive coding is the process of using a smaller set of manual reviewed and coded documents as examples to build a computer generated mathematical model that is then used to predict the coding on a larger set of documents.\u00a0 It is a specialized application of a class of techniques referred to as supervised machine-learning in computer science.\u00a0 Other technical terms often used to describe predictive coding include document (or text) \u201cclassification\u201d and document (or text) \u201ccategorization.\u201d[14]<\/em><\/p>\n<p>\u00a0[11]\u00a0\u00a0\u00a0\u00a0\u00a0 And as stated in <i>The Sedona Conference Best Practices Commentary on the Use of Search and Information Retrieval Methods in E-Discovery <\/i>(Post-Public Comment Version 2013):<\/p>\n<p style=\"padding-left: 30px\"><em>Generally put, computer- or technology-assisted approaches are based on iterative processes where one (or more) attorneys or [Information Retrieval] experts train the software, using document exemplars, to differentiate between relevant and non-relevant documents.\u00a0 In most cases, these technologies are combined with statistical and quality assurance features that assess the quality of the results. \u00a0The research . . . has demonstrated such techniques superior, in most cases, to traditional keyword based search, and, even, in some cases, to human review.<\/em><\/p>\n<p style=\"padding-left: 30px\"><em>\u00a0The computer- or technology-assisted review paradigm is the joint product of human expertise (usually an attorney or IR expert working in concert with case attorneys) and technology.\u00a0 The quality of the application\u2019s output, which is an assessment or ranking of the relevance of each document in the collection, is highly dependent on the quality of the input, that is, the human training. Best practices focus on the utilization of informed, experienced, and reliable individuals training the system.\u00a0 These individuals work in close consultation with the legal team handling the matter, for engineering the application. Similarly . . . the defensibility and usability of computer- or technology-assisted review tools require the application of statistically-valid approaches to selection of a \u201cseed\u201d or \u201ctraining\u201d set of documents, monitoring of the training process, sampling, and quantification and verification of the results.[15]\u00a0<\/em><\/p>\n<p>A discussion of the mathematical algorithms that underlie predictive coding is beyond the intended scope of this Article, but the interested reader should refer to references cited at the margin to understand better what is \u201cgoing on under the hood\u201d with respect to the mathematics involved.[16]<\/p>\n<p>[12]\u00a0\u00a0\u00a0\u00a0\u00a0 <i>The <\/i>Da Silva Moore<i> Precedent.\u00a0 <\/i>The various threads in search and retrieval law, including the need for advanced search methods applied to document review in a world of increasingly large data sets, were well known by 2012.\u00a0 In February 2012, drawing on recent research and scholarship emanating out of the Text Retrieval Conference (TREC) Legal Track[17] and the 2007 public comment version of The Sedona Conference Search Commentary,[18] Judge Peck approached the <i>Da Silva Moore <\/i>case as an appropriate vehicle to provide a judicial blessing for the use of predictive coding in e-Discovery.\u00a0 In doing so, however, Judge Peck\u2019s opinion may also be viewed as setting the stage for greater use of analytics generally in the information governance practice area, beyond \u201cmere\u201d e-Discovery.<\/p>\n<p>[13]\u00a0\u00a0\u00a0\u00a0\u00a0 Plaintiffs in <i>Da Silva Moore<\/i> brought claims of gender discrimination against defendant advertising conglomerate Publicis Groupe and its United States public relations subsidiary, defendant MSL Group.[19]\u00a0 Prior to the February 2012 opinion issued by Judge Peck, the parties had already agreed that defendant MSL would use predictive coding to review and produce relevant documents, but disagreed on methodology.[20]\u00a0 Defendant MSL proposed starting with the manual review of a random sample of documents to create a \u201cseed set\u201d of documents that would be used to train the predictive coding software.[21]\u00a0 Plaintiffs would participate in the creation of the \u201cseed set\u201d of documents by offering keywords.[22] All documents reviewed during the creation of the \u201cseed set,\u201d relevant or irrelevant, would be provided to plaintiffs.[23]<\/p>\n<p>[14]\u00a0\u00a0\u00a0\u00a0\u00a0 After creation of the seed set of documents, MSL proposed using a series of \u201citerative rounds\u201d to test and stabilize the training software.[24]\u00a0 The results of these iterative rounds would be provided to plaintiffs, who would be able to provide feedback to further refine the searches.[25]\u00a0\u00a0 Judge Peck accepted MSL\u2019s proposal.[26]\u00a0 Plaintiffs filed objections with the district judge on the grounds that Judge Peck\u2019s approval of MSL\u2019s protocol unlawfully disposed of MSL\u2019s duty under Federal Rule of Civil Procedure 26(g) to certify the completeness of its document collection, and the methodology in MSL\u2019s protocol was not sufficiently reliable to satisfy Federal Rule of Evidence 702 and <i>Daubert<\/i>.[27]<\/p>\n<p><b>\u00a0<\/b>[15]\u00a0\u00a0\u00a0\u00a0\u00a0 Judge Peck found the plaintiffs\u2019 objections to be misplaced and irrelevant.[28]\u00a0 With respect to Federal Rule of Civil Procedure 26(g), Judge Peck commented that no attorney could certify the completeness of a document production as large as MSL\u2019s. Moreover, Federal Rule of Civil Procedure 26(g) did not require the type of certification plaintiffs described.[29]\u00a0 Further, Federal Rule of Evidence 702 and <i>Daubert<\/i> are applicable to expert methodology, not to methodologies used in electronic discovery.[30]\u00a0 Judge Peck went on to note that the decision to allow computer-assisted review in this case was easy because the parties agreed to this method of document collection and review.[31]\u00a0 While computer-assisted review may not be a perfect system, he found it to be more<br \/>\n efficient and effective than using manual review and keyword searches to locate responsive documents.[32]\u00a0 Use of predictive coding was appropriate in this case considering:<\/p>\n<p style=\"padding-left: 30px\"><em>\u00a0(1) the parties\u2019 agreement, (2) the vast amount of ESI to be reviewed (over three million documents), (3) the superiority of computer-assisted review to the available alternatives (i.e., linear manual review or keyword searches), (4) the need for cost effectiveness and proportionality under Rule 26(b)(2)(C), and (5) the transparent process proposed by MSL.[33]<\/em><\/p>\n<p>\u00a0[16]\u00a0\u00a0\u00a0\u00a0\u00a0 In issuing this opinion, Judge Peck became the first judge to approve the use of computer-assisted review.[34]\u00a0 He also stressed the limitations of his opinion, stating that computer-assisted review may not be appropriate in all cases, and his opinion was not intended to endorse any particular computer-assisted review method.[35]\u00a0 However, Judge Peck encouraged the Bar to consider computer-assisted review as an available tool for \u201clarge-data-volume cases\u201d where use of such methods could save significant amounts of legal fees.[36]\u00a0 Judge Peck also stressed the importance of cooperation, or what he called \u201cstrategic proactive disclosure of information.\u201d\u00a0 If counsel is knowledgeable about the client\u2019s key custodians and fully explains proposed search methods to opposing counsel and the court, those proposed search methods are more likely to be approved.\u00a0 To sum up his opinion, Judge Peck noted that \u201c[c]ounsel no longer have to worry about being the \u2018first\u2019 or \u2018guinea pig\u2019 for judicial acceptance of computer-assisted review. . . . Computer-assisted review now can be considered judicially-approved for use in appropriate cases.\u201d[37]\u00a0 In the two years since <i>Da Silva Moore<\/i>, in addition to cases in which the parties have agreed upon a predictive coding methodology,[38] courts have confronted the issue of having to rule on either the requesting or responding party\u2019s motion to compel a judicial \u201cblessing\u201d of the use of predictive coding (however termed).\u00a0 In <i>Global Aerospace<\/i>,[39] the responding party asked that the court approve its own use of such technique; in <i>Kleen Products, <\/i>the requesting party made an ultimately unsuccessful demand for a \u201cdo-over\u201d in discovery, where the responding party had used keyword search methods and the plaintiffs were demanding that more advanced methods be tried.[40]\u00a0 In the <i>EOHRB <\/i>case, the Court <i>sua sponte <\/i>suggested that the parties consider using predictive coding, including the same vendor.[41]\u00a0 And in the <i>In re Biomet <\/i>case<i>,<\/i>[42]<i> <\/i>the court approved a predictive coding methodology over the objections of the requesting party.\u00a0 These cases represent only some of the reported decisions to date, and we suspect that there will be dozens of reported cases and many more unreported ones in the near term.<\/p>\n<p>[17]\u00a0\u00a0\u00a0\u00a0\u00a0 As recognized in these cases (implicitly or explicitly), as well as in a growing number of commentaries,[43] predictive coding is an analytical technique holding the promise of achieving much greater efficiencies in the e-Discovery process.\u00a0 Notwithstanding <i>Da Silva Moore\u2019s <\/i>call to action, it needs to be conceded, however, that the research has not proven that active machine learning techniques will <i>always<\/i> achieve greater scores than keyword search or manual review.[44] \u00a0Additionally, we bow to the reality that in a large class of cases the use of predictive coding is currently infeasible or unwarranted, especially as a matter of cost.[45]<\/p>\n<p>[18]\u00a0\u00a0\u00a0\u00a0\u00a0 Nevertheless, it seems apparent that the legal profession finds itself in a new place\u2014namely, in need of recognizing that artificial intelligence techniques are growing in strength from year to year\u2014and thus it appears to be only a matter of time until a much greater percentage of complex cases involving a large magnitude of ESI will constitute good candidates for lawyers using predictive coding techniques, both as available currently and as improved with future technological progress.\u00a0 As William Gibson once put it, \u201cthe future is here, it\u2019s just not evenly distributed.\u201d[46]<\/p>\n<h3>\u00a0<b>B.<\/b>\u00a0 <b>Information Governance and Analytics in the Era of Big Data<\/b><\/h3>\n<p>[19]\u00a0\u00a0\u00a0\u00a0\u00a0 We are now in a post-<i>Da Silva Moore<\/i>, \u201cBig data\u201d era where lawyers are on constructive (if not actual) notice of a world of technology assisted review techniques available at least in the sphere of e-Discovery.\u00a0\u00a0The proposition being advanced is that the greater revelation of <i>Da Silva Moore <\/i>is how similar the techniques being put forward as best practices in e-Discovery fit a larger realm of issues familiar to lawyers, many of which fall within what is increasingly being recognized as \u201cinformation governance\u201d practice.\u00a0\u00a0It is here where we can break new ground in our legal practice by recommending the use of these advanced techniques to solve real-world problems of our clients.\u00a0 First, however, some definitions are in order to better frame the legal issues that will follow in Section C.<\/p>\n<p>[20]\u00a0\u00a0\u00a0\u00a0\u00a0 <i>Big data.\u00a0 <\/i>It has been noted that \u201cBig data is a loosely defined term used to describe data sets so large and complex that they become awkward to work with using standard statistical software.\u201d[47]\u00a0 Alternatively, \u201cBig data\u201d is a term that \u201cdescribe[s] the technologies and techniques used to capture and utilize the exponentially increasing streams of data with the goal of bringing enterprise-wide visibility and insights to make rapid critical decisions.\u201d[48]<\/p>\n<p>[21]\u00a0\u00a0\u00a0\u00a0\u00a0 The fact that the data encountered within the corporate enterprise increasingly is indeed \u201cbig\u201d means, at least according to Gartner, that it not only has volume, but velocity and complexity as well.[49] \u00a0As Bill Franks has put it, \u201cWhat this means is that you aren\u2019t just getting a lot of data when you work with big data.\u00a0 It\u2019s also coming at you fast, it\u2019s coming at you in complex formats, and it\u2019s coming at you from a variety of sources.\u201d[50]\u00a0 These elements all significantly contribute to the challenge of finding signals in the noise.<\/p>\n<p>[22]\u00a0\u00a0\u00a0\u00a0\u00a0 These definitions seem to get us closer to what makes Big data a new and interesting phenomenon in the world: it is not its volume alone, but the fact that we are able to \u201cmine\u201d large data sets using new and advanced techniques to uncover unexpected relationships, patterns and categories within these data sets, that makes the field potentially exciting.\u00a0 Indeed, \u201cit is tempting to understand big data solely in terms of size. But that would be misleading. Big data is also characterized by the ability to render into data many aspects of the world that have never been quantified before; call it \u2018datafication.\u2019\u201d[51]<\/p>\n<p>[23]\u00a0\u00a0\u00a0\u00a0\u00a0 <i>Analytics.\u00a0 <\/i>Second, we need to place \u201cpredictive coding\u201d as one form of active machine learning in the context of the broader realm of \u201canalytics.\u201d\u00a0 In their book, <i>Keeping Up With the Quants: Your Guide To Understanding and Using Analytics<\/i>,[52] authors Thomas Davenport and Jinho Kim provide a useful construct in categorizing the newly emergent field of \u201canalytics\u201d: they define analytics to mean \u201cthe extensive use of data, statistical and quantitative analysis, explanatory and predictive models, and fact-based management to drive decisions and add value,\u201d going on to say that \u201c[a]nalytics is all about making sense of big data, and using it for competitive advantage.\u201d\u00a0 The authors divide the world of analytics into three categories:<\/p>\n<p style=\"padding-left: 30px\"><em>\u00a0(i) \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 descriptive analytics \u2013 gathering, organizing, tabulating and depicting data;<\/em><\/p>\n<p style=\"padding-left: 30px\"><em>(ii) \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0 \u00a0predictive analytics \u2013 using data to predict future courses of action; and<\/em><\/p>\n<p style=\"padding-left: 30px\"><em>(iii)\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 prescriptive analytics \u2013 recommendations on future courses of action.[53]<\/em><\/p>\n<p>\u00a0<br \/>\n[24]\u00a0\u00a0\u00a0\u00a0\u00a0 To the extent that \u201cpredictive coding\u201d has been used to date to have machines \u201cpredict\u201d relevancy in large ESI data sets, the term comfortably can be said to fall within category (ii).\u00a0\u00a0 But the world of analytics is a larger universe, encompassing a greater number of mathematical magic tricks,[54] and this should be kept in mind as we choose to limit our discussion here to a few examples of how predictive coding as one form of analytics may be usefully applied in non-traditional contexts.[55]<\/p>\n<p><b>\u00a0<\/b>[25]\u00a0\u00a0\u00a0\u00a0\u00a0 Corporations (much ahead of the legal profession) have rushed headlong during the past half-decade to use a variety of analytics to understand the Big data they increasingly hold, to add value, and to improve the bottom line.[56]\u00a0 A 2013 AIIM study indicates that corporations find analytics to be useful in a variety of settings.[57]<\/p>\n<p>[26]\u00a0\u00a0\u00a0\u00a0\u00a0 <i>Information Governance.\u00a0 <\/i>\u201cInformation governance,\u201d as defined in The Sedona Conference\u2019s recently published Commentary on the subject, means:<\/p>\n<p style=\"padding-left: 30px\"><em>\u00a0an organization\u2019s coordinated, interdisciplinary approach to satisfying information legal and compliance requirements and managing information risks while optimizing information value.\u00a0 As such, Information Governance encompasses and reconciles the various legal and compliance requirements and risks addressed by different information focused disciplines, such as records and information management (\u201cRIM\u201d), data privacy, information security, and e-[D]iscovery.[58]<\/em><\/p>\n<p>\u00a0Or, as highlighted by the seminal law review article devoted to information governance written by Charles R. Ragan who quotes Barclay Blair in defining information governance as a \u201c\u2018new approach\u2019 that \u201cbuilds upon and adapts disciplines like records management and retention, archiving business analytics, and IT governance to create an integrated model for harnessing and controlling enterprise information . . . [I]t is an evolutionary model that requires organizations to make real changes.\u201d[59]<\/p>\n<p>[27]\u00a0\u00a0\u00a0\u00a0\u00a0 As the Sedona IG Commentary highlights, \u201cmany organizations have traditionally used siloed approaches when managing information.\u201d[60]\u00a0 The \u201ccore shortcoming\u201d of this approach is \u201cthat those within particular silos are constrained by the culture, knowledge, and short-term goals of their business unit, administrative function, or discipline.\u201d[61]\u00a0 This leads in turn to key actors within the organization having \u201cno knowledge of gaps and overlaps in technology or information in relation to other silos. . . .\u201d[62]\u00a0 In such situations, \u201c[t]here is no overall governance or coordination for managing information as an asset, and there is no roadmap for the current and future use of information technology.\u201d[63]<\/p>\n<p>[28]\u00a0\u00a0\u00a0\u00a0\u00a0 The Sedona IG Commentary goes on to provide eleven principles of what constitutes good IG practices, of which Principle 10 is of special relevance to our discussion here: \u201cAn organization should consider leveraging the power of new technologies in its Information Governance program.\u201d[64]\u00a0 As stated therein,<\/p>\n<p style=\"padding-left: 30px\">\u00a0\u00a0 \u00a0 \u00a0 \u00a0 <em>Organizations should consider using advanced tools and technologies to perform various types of categorization and classification activities. . . such as machine learning, auto-categorization, and predictive analytics to perform multiple purposes, including (i) optimizing the governance of information for traditional RIM [records and information management]; (ii) providing more efficient and more efficacious means of accessing\u00a0 information for e-discovery, compliance, and open records laws, and (iii) advancing sophisticated business intelligence across the enterprise.[65]<\/em><\/p>\n<p><i>\u00a0<\/i>With respect to the latter category, the Commentary goes on to specifically identify areas where predictive analytics may be used in compliance programs \u201cto predict and prevent wrongful or negligent conduct that might result in data breach or loss,\u201d as a type of \u201cearly warning system.\u201d[66] It is precisely this latter type of conduct that we wish to primarily explore in the next section, along with a few final words on using analytics with auto-categorization for the purpose of records classification and data remediation.<\/p>\n<h3><b>C.\u00a0 Applying the Lessons of E-Discovery In Using Analytics for Optimal Information Governance: Some Examples<\/b><\/h3>\n<p><b>\u00a0<\/b>[29]\u00a0\u00a0\u00a0\u00a0\u00a0 Advanced analytics are increasingly being used in the e-Discovery context because the legal profession has begun to realize the limitations of manual and keyword searching, while at the same time seeing how advanced techniques are at least as efficacious and far more efficient in a wide variety of substantial engagements.\u00a0 But more efficient and at least as equally effective at doing what, precisely?\u00a0 In e-Discovery, the primary information task involves separating relevant from non-relevant, and to a secondary degree, privileged from non-privileged information, in documents and ESI.\u00a0 Indeed, lawyers are under a duty to make \u201creasonable\u201d\u2014not perfect\u2014efforts to find <i>all <\/i>relevant documents within the scope of a given discovery request.[67]\u00a0 The illusiveness of this quest in an exponentially expanding data universe is becoming increasingly apparent to many.[68]<\/p>\n<p>[30]\u00a0\u00a0\u00a0\u00a0\u00a0 Moreover, the degree of success in being able to either find or demand substantial amounts of relevant information is not (nor should it be) the fundamental goal or point of engaging in e-Discovery.[69]\u00a0 Rather, the liberal discovery rules that at least U.S. lawyers operate within have as their underlying purpose the ferreting out of important, material facts to the case at hand.\u00a0 The increasingly overwhelming nature of ESI poses clear technological obstacles to a lawyer en route to efficiently engaging in developing facts from all those relevant documents to determine what happened and why.[70]\u00a0 The promise of using an advanced analytical method such as predictive coding is its ability to quickly find and rank-order the <i>most <\/i>relevant documents for answering these questions.\u00a0 For once we determine how something happened and why, it is relatively straightforward to figure out the parties\u2019 respective rights, responsibilities, and even liability.\u00a0 That is precisely the point of litigation, and the purpose of the Rules that govern it.[71]\u00a0 And, facts drive it all.<\/p>\n<p>[31]\u00a0\u00a0\u00a0\u00a0\u00a0 Given our increasing ability in litigation in finding the most relevant needles (i.e., facts) in the Big data haystack, it stands to consider whether similar methods may be successfully applied in non-litigation contexts.\u00a0 Somewhat paradoxically, however, experience indicates that there are advantages to dealing with <i>larger <\/i>volumes of data when applying analytical tools and methods to solve corporate legal issues.\u00a0 That is, while a vast amount of data residing in corporate networks and repositories admittedly poses complex information governance challenges, the volume of Big data also may be a boon to the investigator simply trying to figure out what happened.\u00a0 This is the case because there are simply many more data points from which to derive facts.\u00a0 One can liken the phenomenon to the difference in quality of a one-megapixel versus a ten-megapixel picture: the difference in the quality of the image is a function of the greater density of points of illumination.<\/p>\n<p>[32]\u00a0\u00a0\u00a0\u00a0\u00a0 Big data is more data, and more data means the potential for a more complete picture of what happened in a given situation of interest, assuming of course that the facts can be captured <i>efficiently<\/i>.\u00a0 The problem is not one of volume, but of visibility.\u00a0 In the era of Big data, the investigator with the more powerful analytical methods, who can search into vast repositories of ESI to draw out the facts that are critical to the question at hand, is king (or queen).\u00a0 This is where the skillful application of advanced analytics to Big data can bring about some remarkable results.\u00a0 The true strategic advant<br \/>\nage of advanced analytics is the <i>speed<\/i> with which an accurate answer can be ascertained.[72]<\/p>\n<p>[33]\u00a0\u00a0\u00a0\u00a0\u00a0 <i>True Life Example #1.<\/i>[73]\u00a0 A corporate client is being sued by a former employee in a whistleblower <i>qui tam <\/i>action.[74]\u00a0 Because of the False Claims Act allegations, the suit represented a significant threat to the company.\u00a0 The corporation retains counsel to understand the client\u2019s information systems as well as its key players, and to assist in the implementation of a litigation hold.\u00a0 Counsel strategically targets the data most likely to shed light on the facts.\u00a0 The law firm\u2019s Fact Development Team applies advanced analytics to 675,000 documents, and within four days knows enough to defend the client\u2019s position that the allegations are indisputably baseless.\u00a0 All of this is done before the answer to the Complaint was due.<\/p>\n<p>[34]\u00a0\u00a0\u00a0\u00a0\u00a0 Armed with this information, counsel for the corporation approached plaintiff\u2019s counsel and asked to meet.\u00a0 Prior to the meeting, the corporation voluntarily produced 12,500 documents that laid out the parties\u2019 position precisely.\u00a0 Counsel then met with plaintiff\u2019s counsel and walked them through the evidence, laying out all the facts.\u00a0 The case ended up being settled within days for what amounted to nuisance value based on a retaliation claim\u2014without any discovery, and at a small fraction of the cost budgeted for the litigation.<\/p>\n<p>[35]\u00a0\u00a0\u00a0\u00a0\u00a0 This example indicates that the real power of advanced analytics is not merely in potentially reducing the cost of vexatious litigation, but rather the strategic<i> advantage<\/i> that comes with counsel getting to an answer quickly and accurately.\u00a0 This precise strategic advantage has many applications outside of litigation, each of which involves an aspect of optimizing information governance.<\/p>\n<p>[36]\u00a0\u00a0\u00a0\u00a0\u00a0 Only a short step away from the direct litigation realm is using advanced analytics for investigations, either in response to a regulatory inquiry or for purely internal purposes.\u00a0 As we have already seen, corporate clients are often faced with circumstances where determining whether an allegation is true, and the scope of the potential problem if it is, is critically important.\u00a0 Often, management must wait, unsure of their company\u2019s exposure and how to remediate it, while traditional investigation techniques crawl along.\u00a0 However, with the skillful application of advanced analytics upon the right data set, accurate answers can be determined with remarkable speed.<\/p>\n<p>[37]\u00a0\u00a0\u00a0\u00a0\u00a0 <i>True Life Example #2.<\/i>\u00a0 A highly regulated manufacturing client decided to outsource the function of safety testing some of its products.\u00a0 A director of the department whose function was being outsourced was offered a generous severance package.\u00a0 Late on a Friday afternoon, the soon-to-be former director sent an email to the company\u2019s CEO demanding four times the severance amount and threatened to go to the company\u2019s regulator with a list of ten supposed major violations that he described in the email if he did not receive what he was asking for.\u00a0 He gave the company until the following Monday to respond.<\/p>\n<p>[38]\u00a0\u00a0\u00a0\u00a0\u00a0 The lawyers were called in.\u00a0 They analyzed the list of allegations and determined which IT systems would most likely contain data that would prove their veracity and immediately pulled the data.\u00a0 Applying advanced analytics, the law firm\u2019s Fact Development Team analyzed on the order of 275,000 documents in thirty-six hours.\u00a0 By that Monday morning, counsel was able to present a report to the company\u2019s board indisputably proving that the allegations were unfounded.<\/p>\n<p>[39]\u00a0\u00a0\u00a0\u00a0\u00a0 <i>True Life Example #3<\/i>.\u00a0 A major company received a whistleblower letter from a reputable third party alleging that several senior personnel were involved with an elaborate kickback scheme that also involved FCPA violations.\u00a0 If true, the company would have faced serious regulatory and legal issues, as well as major internal difficulties.\u00a0 Because of the extremely sensitive nature of the allegations, a traditional investigation was not possible; even knowing certain personnel were under investigation could have had immense consequences.<\/p>\n<p>[40]\u00a0\u00a0\u00a0\u00a0\u00a0 The lawyers were tasked with determining whether there was any information within the company\u2019s possession that shed any light on the allegations.\u00a0 If there were, the company would proceed to take whatever steps were required.\u00a0 The investigation was of such a secret nature that no one was authorized to involve the internal IT staff.\u00a0 Fortunately, counsel knew the company and its information systems well.\u00a0 Over a weekend, they were able to pull 8.5 million documents from relevant systems using the law firm\u2019s personnel.\u00a0 This turned out to be a highly complex investigation involving a number of potential subjects, where the task involved tracking the subject\u2019s travel, meetings with suppliers, subsequent sales orders and fulfillments, rebates and promotions, all across several years.<\/p>\n<p>[41]\u00a0\u00a0\u00a0\u00a0\u00a0 Again, applying advanced analytics, the law firm\u2019s Fact Development Team analyzed the 8.5 million documents in ten days.\u00a0 They were able to prove that the allegations were largely baseless, and precisely where there were potential areas of concern.\u00a0 Counsel also was able to make clear recommendations for areas of further investigation and for modifying compliance tracking and programs.\u00a0 The company was able to act quickly and with certainty.\u00a0 These real-life use cases illustrate how the power of analytics enhances the ability of lawyers to provide legal advice under conditions of \u201ccertainty\u201d previously unobtainable, at least in the past few decades of the digital era.\u00a0 \u201cCertainty\u201d is a somewhat foreign concept in the law\u2014lawyers tend to be a conservative and caveating bunch, largely because certainty has historically been hard to come by, or at least prohibitively expensive.\u00a0 With advanced analytics and good lawyers who know how to use these new tools, that is no longer necessarily the case.\u00a0 There is so much data that if one cannot, after a reasonable effort, find evidence of a fact in the vastness of a company\u2019s electronic information (as long as you have the right information), the fact most likely is not true.\u00a0 Such has been illustrated, proving a negative is particularly useful in investigations.<\/p>\n<p>[42]\u00a0\u00a0\u00a0\u00a0\u00a0 Using advanced analytics (and good lawyering) for investigations is not that far removed from using it for litigation: one is still attempting to find the answer to the question of what happened and why. But there are many other questions that companies would like to ask of their data.\u00a0 And indeed, both the analytics tools and the fact development techniques used in litigation and investigations can be \u201ctuned\u201d to solve a variety of novel issues facing our clients.<\/p>\n<p>[43]\u00a0\u00a0\u00a0\u00a0\u00a0 For example, analytics can be used to vet candidates for political appointments as well as candidates for senior leadership positions.\u00a0 Due to the candid nature of the medium, providing access to corporate email coupled with using analytic capabilities allows for an accurate picture to be drawn <i>before<\/i> a decision is made with regard to making a candidate your next CEO or running mate.\u00a0 Analytics can be used to analyze business divisions to identify good and bad leaders, how decisions are made, why a division is more successful than another, and many more similar applications.<\/p>\n<p>[44]\u00a0\u00a0\u00a0\u00a0\u00a0 Quite simply, a company\u2019s data is the digital imprint of the actions and decisions of all of its managers and employees.\u00a0 Having insight into those actions and decisions can be immensely valuable.\u00a0 That value has lain largely fallow, hidden in plain sight because the valuable wheat could not effectively be sifted from the chaff.\u00a0 With the proper application of advanced analytics, that is no longer the case.\u00a0 The answers we can obtain are limited only by the creativity of management in asking the right questions.<\/p>\n<p>[45]\u00a0\u00a0\u00a0\u00a0\u00a0 <i>True Life Example #4.<\/i>\u00a0 Advanced a<br \/>\nnalytics used upon the major acquisition of another company by a corporate client.\u00a0 As with most acquisitions, the client undertook traditional due diligence, gathering information from the target regarding its financial performance, customers, market share, receivables, potential liabilities, and came up with a valuation, an appropriate multiplier, and a final purchase price.\u00a0 Also as is typical, the acquisition agreement contained a provision such that if the disclosures made by the target were found to be off by a certain margin within thirty days of the acquisition, the purchase price would be adjusted.<\/p>\n<p>[46]\u00a0\u00a0\u00a0\u00a0\u00a0 The moment the acquisition closed, the corporate client then owned all of the target\u2019s information systems.\u00a0 Having some concern about the bases for some of the target\u2019s disclosures, at the client\u2019s request counsel proceeded to use analytics on those newly acquired systems to determine what we could about those disclosures.\u00a0 Preparing a company for sale is a complicated affair, with many people involved in gathering information to present to the acquirer to satisfy due diligence.\u00a0 This gathering and presentation of information is done primarily through electronic means\u2014and leaves a trail.<\/p>\n<p>[47]\u00a0\u00a0\u00a0\u00a0\u00a0 Using advanced analytics, the law firm\u2019s Fact Development Team traced the compilation of the target\u2019s due diligence information, including all of the discussion that went along with it.\u00a0 They were able to understand the source of each disclosure, the reasonableness of its basis, and any weaknesses within it.\u00a0 They uncovered disagreements within the target over such things as what the right numbers were, or how much of a liability to disclose.\u00a0 Using this information, counsel prepared a claim in accord with the adjustment provision seeking twenty-five percent of the purchase price totaling millions of dollars.\u00a0 The claim was primarily composed using quotes from their own documents.\u00a0 It is difficult to argue with yourself.<\/p>\n<p>[48]\u00a0\u00a0\u00a0\u00a0\u00a0 As demonstrated, using advanced analytics in the form of predictive coding and similar technologies can accomplish some notable aims.\u00a0 But each of the prior examples uses data to look back to determine what has already occurred: the descriptive use of analytics.[75]\u00a0 This is extremely valuable.\u00a0 But for many of a law firm\u2019s clients, it would be even more useful to be able to catch bad actors while the misconduct was occurring, or even to predict misconduct before it happens.<\/p>\n<p>[49]\u00a0\u00a0\u00a0\u00a0\u00a0 Based on the anecdotal experience gathered from many past investigations, the authors believe that certain kinds of misconduct follow certain patterns, and that when bad actors are acting badly, they tend to undertake the same kinds of actions, or are experiencing similar circumstances.\u00a0 For example, in our experience the primary factors that pertain to a person committing fraud are personal relationship problems, financial difficulties, drug or alcohol problems, gambling, a feeling of under appreciation at work, and unreasonable pressure to achieve a work outcome without a legitimate way to accomplish it (and so they attempt illegitimate ways to do so).\u00a0 These factors are often detectable in the electronic information the subject creates.\u00a0 Similarly, a person who is harassing or discriminating against others also tends to undertake specific actions and use particular language in communications.\u00a0 All of these indicia of misconduct are detectable using advanced analytics and skillful strategy.<\/p>\n<p>[50]\u00a0\u00a0\u00a0\u00a0\u00a0 Lawyers have gotten quite good at finding this information when looking back in time.\u00a0 We thought, then, that it should not be too difficult to find this information while the misconduct is unfolding, or to identify warning signs that misconduct is likely to occur, and seek to provide relief of certain factors where possible or take corrective action when needed and as early as possible.\u00a0 So, we put this to the test, developing Early Warning Systems (\u201cEWS\u201d) for some of our clients.<\/p>\n<p>[51]\u00a0\u00a0\u00a0\u00a0\u00a0 The idea for an EWS first occurred to one of the authors when working on a pro bono matter with the ACLU in a case against the Baltimore Police Department (\u201cBPD\u201d) alleging unconstitutional arrest practices in its Zero Tolerance Policing policies.[76]\u00a0 As a result of the case, the BPD agreed to, among other things, implement a tracking system whereby certain data points were collected regarding police officer conduct and arrest practices that research had proven were warning signs of potential problem officers.[77]\u00a0 The accumulation of certain data points with respect to an officer triggered a review of the officer\u2019s conduct, with various remediation outcomes.[78]\u00a0 We thought that a similar approach could be used for our clients.<\/p>\n<p>[52]\u00a0\u00a0\u00a0\u00a0\u00a0 An EWS is a tricky thing to implement, and requires careful consideration of many factors, employee privacy at the forefront.\u00a0 However, with careful planning, policy development, and training, an effective EWS can be designed and implemented.\u00a0 Predictive analytics applications can be trained to search for indicia of the conduct, language, or factors across information systems.\u00a0 The specific systems to be targeted will vary depending on what is being sought and the systems most likely to contain it and will vary greatly from company to company.\u00a0 But, when properly trained and targeted, we have found these systems to be very effective in detecting and even preventing misconduct.\u00a0 We believe that this use of predictive analytics will become one of the most powerful applications of this technology in the near future.<\/p>\n<p>[53]\u00a0\u00a0\u00a0\u00a0\u00a0 Moving from the business intelligence aspects of information governance to the arguably more prosaic field of records and information management, the authors also count themselves as true believers in the power of analytics to optimize traditional RIM (records and information management) functionality.\u00a0 A full discussion of archival and records management practices in the digital age is beyond the scope of this Article, but the interested reader will find a wealth of scholarly literature in the leading journals discussing how the traditional practice of records management is being transformed in the digital age. One of the authors has argued that predictive coding and like methods are the most promising way to open up \u201cdark archives\u201d in the public sector, such as digital collections of data appraised as permanent records (mostly consisting of White House email at this point), that for reasons of privacy or privilege will be otherwise inaccessible to the public for many decades to come.[79]<\/p>\n<p>[54]\u00a0\u00a0\u00a0\u00a0\u00a0 In the authors\u2019 experience, email archiving using auto-categorization for recordkeeping purposes is available using existing software in the marketplace.\u00a0 In such instances, email is populated in specific \u201cbuckets\u201d in a repository depending on how it is characterized, based on either the position of the creator or recipient of the email, the subject matter, or based on some other attribute appearing as metadata.[80]\u00a0 In the most advanced versions of auto-categorization software, the system \u201clearns\u201d as it is trained using exemplars in a seed set selected by subject matter experts (i.e., records managers or expert end users), via a protocol highly reminiscent of the methods adopted by the parties in <i>Da Silva Moore <\/i>and similar cases.\u00a0 It is only a matter of time before predictive analytics is more widely used to optimize auto-classification while reducing the burden on end users to perform manual records management functions.[81]<\/p>\n<p>[55]\u00a0\u00a0\u00a0\u00a0\u00a0 In similar fashion, the power of predictive analytics to reliably classify content after adequate training makes such tools optimal for data remediation efforts.\u00a0 The problem of legacy data in corporations is well known, and only growing over time with the inflationary expansion of the ESI universe.[82]\u00a0 Using advanced analytics to classify low value data, the chaos that is the reality of most shared drives and other joint data repositories, may potentially be reduced by<br \/>\norders of magnitude.\u00a0 The challenge of engaging in defensible deletion is one important aspect of optimizing information governance.[83]<\/p>\n<p><b>\u00a0<\/b><\/p>\n<h2 align=\"center\"><b>Conclusion<\/b><\/h2>\n<p><b>\u00a0<\/b>[56]\u00a0\u00a0\u00a0\u00a0\u00a0 As was made clear at the outset, it is the authors\u2019 intent merely to scratch the surface of what is possible in the analytics space as applied to matters of importance for corporate information governance.\u00a0 No one has a one hundred percent reliable crystal ball, but it seems evident that as computing power increases, those forms of artificial intelligence that we have referred to here as analytics will themselves only grow in importance in both our daily and professional lives.\u00a0 By the end of this decade, we would be surprised if the following do <i>not <\/i>occur: pervasive use of business intelligence software; the use of more automated decision-making (also known as \u201coperational business intelligence\u201d); the use of alerts in the form of early warning systems including the type described above; much greater use of text mining and predictive technologies across a variety of domains.[84]<\/p>\n<p>[57]\u00a0\u00a0\u00a0\u00a0\u00a0 All of these developments dovetail with the expected demand on the part of corporate clients for lawyers to be familiar with state of the art practices in the information governance space, as already anticipated by the type of technology that <i>Da Silva Moore <\/i>and related cases suggest.\u00a0 As best said in <i>The Sedona Commentary on Achieving Quality in E-Discovery<\/i>, \u201c[i]n the end, cost-conscious firms, organizations, and institutions of all types that are intent on best practices . . . will demand that parties undertake new ways of thinking about how to solve e-[D]iscovery problems. . . .\u201d [85]\u00a0 The same holds true for the greater playing field of information governance.\u00a0 Lawyers who have embraced analytics will have a leg up on their competition in this brave new space.<\/p>\n<p>&nbsp;<\/p>\n<div>\n<hr align=\"left\" size=\"1\" width=\"33%\" \/>\n<div>\n<p>* Mr. Borden is a partner in the Commercial Litigation section at Drinker Biddle &amp; Reath, LLP, Washington, D.C., where he serves as Chair of the Information Governance and e-Discovery Group.\u00a0 He is Co-Chair of the Cloud Computing Committee and Vice Chair of the e-Discovery and Digital Evidence Committee of the Science and Technology Law Section of the ABA.\u00a0 He is also a founding member of the steering committee for the Electronic Discovery Section of the District of Columbia Bar.\u00a0 B.A., with highest honors, George Mason University; J.D., <i>cum laude<\/i>, Georgetown University Law School.<\/p>\n<p>**\u00a0Mr. Baron serves as Of Counsel in the Information Governance and e-Discovery Group, Drinker Biddle &amp; Reath, LLP, Washington, D.C, and is on the Adjunct Faculty at the University of Maryland.\u00a0 He formerly served as Director of Litigation at the National Archives and Records Administration, and is a former steering committee Co-Chair of The Sedona Conference Working Group 1 on Electronic Document Retention and Production.\u00a0 B.A.,\u00a0<i>magna cum laude,\u00a0<\/i>Wesleyan University; J.D., Boston University School of Law.\u00a0 The authors wish to thank Drinker Biddle &amp; Reath associates Amy Frenzen and Nicholas Feltham for their assistance in the drafting of this article.\u00a0 The views expressed are the authors\u2019 own and do not necessarily reflect the views of any institution, public or private, that they are affiliated with.<\/p>\n<p>&nbsp;<\/p>\n<\/div>\n<div>\n<p>[1] <i>See infra<\/i> text accompanying notes 47-49 for a definition.<\/p>\n<p>[2] Da Silva Moore v. Publicis Groupe, 287 F.R.D. 182, 192 (S.D.N.Y. 2012), <i>aff\u2019d<\/i><i>sub nom<\/i>. Moore v. Publicis Groupe SA,<i> <\/i>2012 U.S. Dist LEXIS 58742 (S.D.N.Y. Apr. 26, 2012) (Carter, J.).<\/p>\n<div>\n<p>[3] We recognize the paradox of articles living \u201cforever\u201d on the Internet, especially when published in online journals such as this one, while at the same time ever more rapidly becoming obsolete and out of date.\u00a0<\/p>\n<\/div>\n<div>\n<p>[4]<i>See generally<\/i> The Sedona Conference, <i>The Sedona Conference Best Practices Commentary on the Use of Search and Information Retrieval Methods in E-Discovery, <\/i>8 Sedona Conf. J. 189, 198 (2007) [hereinafter <i>Sedona Search Commentary<\/i>].<\/p>\n<\/div>\n<div>\n<p>[5] Thomas H. Davenport &amp; Jinho Kim, Keeping Up with the Quants: Your Guide To Understanding and Using Analytics 1-2 (2013).<\/p>\n<\/div>\n<div>\n<p>[6] <i>See Sedona Search Commentary<\/i>,<i> supra <\/i>note 4, at 200-201 nn.16-19.<\/p>\n<p>[7]<i>See, e.g.<\/i>, George L. Paul &amp; Jason R. Baron, <i>Information Inflation: Can The Legal System Adapt?, <\/i>13 Rich. J.L. &amp; Tech. 10, \u00b6 2 (2007), http:\/\/law.richmond.edu\/jolt\/v13i3\/article10.pdf.<\/p>\n<\/div>\n<div>\n<p>[8]<i> Id.<\/i>; <i>see Sedona Search Commentary<\/i>, <i>supra<\/i> note 4, at 201-202; Mia Mazza, Emmalena K. Quesada, &amp; Ashley L. Stenberg, <i>In Pursuit of FRCP1: Creative Approaches to Cutting and Shifting Costs of Discovery of Electronically Stored Information<\/i>, 13 Rich. J.L. &amp; Tech. 11, \u00b6 46 (2007), http:\/\/jolt.richmond.edu\/v13i3\/article11.pdf.<\/p>\n<\/div>\n<div>\n<p>[9]<i>See<\/i> Victor Stanley v. Creative Pipe, 250 F.R.D. 251, 256-7 (D. Md. 2008); <i>see also<\/i> United States v. O\u2019Keefe, 537 F. Supp. 2d 14, 23-24 (D.D.C. 2008); William A. Gross Const. Ass\u2019n v Am. Mfrs. Mut. Ins. Co., 256 F.R.D. 134, 135 (S.D.N.Y. 2009); Equity Analytics, LLC v. Lundin, 248 F.R.D. 331, 333 (D.D.C. 2008); <i>In re<\/i> Seroquel Prod. Liab. Litig., 244 F.R.D. 650, 663 (M.D. Fla. 2007).\u00a0 <i>See generally<\/i> Jason R. Baron, <i>Law in the Age of Exabytes: Some Further Thoughts on \u2018Information Inflation\u2019 and Current Issues in E-Discovery Search<\/i>,<i> <\/i>17 Rich. J.L. &amp; Tech. 9, \u00b6 11 n.38 (2011),<i> <\/i>http:\/\/jolt.richmond.edu\/v17i3\/article9.pdf.<\/p>\n<\/div>\n<div>\n<p>[10]<i> See <\/i>Paul &amp; Baron, <i>supra <\/i>note 7, at \u00b6 20; <i>see also<\/i> Bennett B. Borden, <i>The Demise of Linear Review<\/i>, Williams Mullen E-Discovery Alert, Oct. 2010, at 1, http:\/\/www.clearwellsystems.com\/e-discovery-blog\/wp-content\/uploads\/2010\/12\/E-Discovery_10-05-2010_Linear-Review_1.pdf.<\/p>\n<\/div>\n<div>\n<p>[11] The Sedona Conference, The Sedona Conference Commentary on Achieving Quality in e-Discovery 1 (Post-Public Comment Version 2013), <i>available at <\/i>www.thesedonaconference.org\/publications (for publication 15 Sedona Conf. J. ___ (2014) (forthcoming)).<\/p>\n<\/div>\n<div>\n<p>[12]<i>Id.<\/i><\/p>\n<\/div>\n<div>\n<p>[13] <i>See, e.g.<\/i>, <i>William A. Gross Constr.<\/i>, 256 F.R.D. at 136; <i>Seroquel<\/i>, 244 F.R.D. at 662.\u00a0 <i>See generally<\/i> Bennett B. Borden et al., <i>Four Years Later: How the 2006 Amendments to the Federal Rules Have Reshaped the E-Discovery Landscape and Are Revitalizing the Civil Justice System<\/i>, 17 Rich. J.L. &amp; Tech. 10, \u00b6\u00b6 30-37 (2011), http:\/\/jolt.richmond.edu\/v17i3\/article10.pdf; Ralph C. Losey, <i>Predictive Coding and the Proportionality Doctrine: A Marriage Made in Big Data<\/i>, 26 Regent U. L. Rev. 7, 53 n.189 (2013) (collecting cases on proportionality).<\/p>\n<\/div>\n<div>\n<p>[14] Rajiv Maheshwari, Predictive Coding Guru\u2019s Guide 21 (2013); <i>see also<\/i> Baron, <i>supra <\/i>note 9, at \u00b6 32, n.124 (stating predictive coding and other like terminology as used by e-Discovery vendors); Maura R. Grossman &amp; Gordon V. Cormack, <i>The Grossman-Cormack Glossary of Technology-Assisted Review<\/i>, 7 Fed. Cts. L. Rev. 1, 4 (2013), http:\/\/www.fclr.org\/fclr\/articles\/html\/2010\/grossman.pdf; Nicholas M. Pace &amp; Laura Zakaras, <i>Where the Money Goes: Understanding Litigant Expenditures for Producing Electronic Discovery<\/i>, RAND Institute for Civil Justice 59 (2012), <i>available at<\/i> http:\/\/www.rand.org\/pubs\/monographs\/MG1208.html (defining predictive coding).<\/p>\n<\/div>\n<div>\n<p>[15] The Sedona Conference, The Sedona Conference Best Practices Commentary on the Use of Search and Information Retrieval Methods in E-Discovery (Post-Public Comment Version 2013), <i>available at<\/i> www.thesedonaconference.org\/publications (for publication in 15 Sedona Conf. J. ___ (2014)).\u00a0 For an excellent, in-depth discussion of how a practitioner may use predictive coding in e-Di<br \/>\nscovery, with references to experiments by the author, see Losey, <i>supra <\/i>note 13, at 9.\u00a0<\/p>\n<\/div>\n<div>\n<p>[16] <i>See, e.g.<\/i>, <i>Sedona Search Commentary<\/i>, <i>supra<\/i> note 4, at app. 217-223 (describing various search methods); Douglas W. Oard &amp; William Webber, <i>Information Retrieval for E-Discovery<\/i>, 7 Foundations and Trends in Information Retrieval 100 (2013), <i>available at <\/i>http:\/\/terpconnect.umd.edu\/~oard\/pdf\/fntir13.pdf; Jason R. Baron &amp; Jesse B. Freeman, <i>Cooperation, Transparency, and the Rise of Support Vector Machines in E-Discovery: Issues Raised By the Need to Classify Documents as Either Responsive or Nonresponsive<\/i> (2013), http:\/\/www.umiacs.umd.edu\/~oard\/desi5\/additional\/Baron-Jason-final.pdf.\u00a0 For good resources in the form of information retrieval textbooks, see Gary Miner, et al., Practical Text Mining and Statistical Structured Text Data Applications (Elsevier: Amsterdam) (2012); Christopher D. Manning, Prabhakar Raghavan, &amp; Hinrich Schutze, Introduction to Information Retrieval\u00a0 (2008).<\/p>\n<\/div>\n<div>\n<p>[17] <i>See<\/i><i>TREC Legal Track<\/i>, U. Md., http:\/\/trec-legal.umiacs.umd.edu (last visited Feb. 23, 2014) (collecting Overview reports from 2006-2011) (as explained on its home page, \u201c[t]he goal of the Legal Track at the Text Retrieval Conference (TREC) [was] to assess the ability of information retrieval techniques to meet the needs of the legal profession for tools and methods capable of helping with the retrieval of electronic business records, principally for use as evidence in civil litigation.\u201d);<b> <\/b><i>see also<\/i> Maura R. Grossman &amp; Gordon V. Cormack, <i>Technology-Assisted Review in E-Discovery Can Be More Effective and More Efficient than Exhaustive Manual Review<\/i>, 17 Rich. J.L. &amp; Tech. 11, \u00b6\u00b6 3-4 (2011), http:\/jolt.richmond.edu\/v17i3\/article11.pdf; Patrick Oot, et al., <i>Mandating Reasonableness in a Reasonable Inquiry, <\/i>87 Denv. U.L. Rev. 533, 558-559 (2010); Herbert Roitblat et al., <i>Document Categorization in Legal Electronic Discovery: Computer Classification vs. Manual Review, <\/i>61 J. Am. Soc\u2019y for Info. Sci. &amp; Tech. 70, 77-79 (2010), <i>available at<\/i> http:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/asi.21233\/full; <i>see generally<\/i> Pace &amp; Zakaras, <i>supra <\/i>note 14<i>, <\/i>at 77-80.<\/p>\n<\/div>\n<div>\n<p>[18] <i>Sedona Search Commentary<\/i>, <i>supra<\/i> note 4, at 192-193.<\/p>\n<\/div>\n<div>\n<p>[19] Da Silva Moore v. Publicis Groupe, 287 F.R.D. 182, 183 (S.D.N.Y 2012), <i>aff\u2019d<\/i> <i>sub nom<\/i>. Moore v. Publicis Groupe SA,<i> <\/i>2012 U.S. Dist LEXIS 58742 (S.D.N.Y. Apr. 26, 2012) (Carter, J.).<\/p>\n<p>[20] <i>Id.<\/i> at 184-87.<\/p>\n<\/div>\n<div>\n<p>[21] <i>Id.<\/i> at 186-87.<\/p>\n<\/div>\n<div>\n<p>[22] <i>Id.<\/i> at 187.<\/p>\n<\/div>\n<div>\n<p>[23] <i>Id.<\/i><\/p>\n<\/div>\n<div>\n<p>[24] <i>Da Silva Moore<\/i>, 287 F.R.D. at 187.<\/p>\n<\/div>\n<div>\n<p>[25] <i>Id.<\/i><\/p>\n<\/div>\n<div>\n<p>[26] <i>Id.<\/i><\/p>\n<\/div>\n<div>\n<p>[27] <i>Id.<\/i> at 188-89.<\/p>\n<\/div>\n<div>\n<p>[28] <i>Id.<\/i><\/p>\n<\/div>\n<div>\n<p>[29] <i>Da Silva Moore<\/i>, 287 F.R.D.<i> <\/i>at 188.<\/p>\n<\/div>\n<div>\n<p>[30] <i>Id.<\/i> at 188-89 (citing Daubert v. Merrell Dow Pharms., 509 U.S. 579, 585 (1993)).\u00a0 <i>But cf. <\/i>David J. Waxse &amp; Benda Yoakum-Kris, <i>Experts on Computer-Assisted Review: Why Federal Rule of Evidence 702 Should Apply to Their Use<\/i>, 52 Washburn L.J. 207, 219-23 (2013) (arguing that the <i>Daubert <\/i>standard should be applied to experts presenting evidence on ESI search and review methodologies)<\/p>\n<\/div>\n<div>\n<p>[31] <i>Id. <\/i>at 189.<\/p>\n<\/div>\n<div>\n<p>[32] <i>Id.<\/i> at 190-91; <i>see <\/i>Grossman &amp; Cormack,<i> supra <\/i>note 17, at \u00b6 61.<\/p>\n<\/div>\n<div>\n<p>[33] <i>Da Silva Moore<\/i>, 287 F.R.D. at 192.<\/p>\n<\/div>\n<div>\n<p>[34] <i>Id.<\/i> at 193.<\/p>\n<\/div>\n<div>\n<p>[35] <i>Id.<\/i><\/p>\n<\/div>\n<div>\n<p>[36] <i>Id.<\/i><\/p>\n<\/div>\n<div>\n<p>[37] <i>Id.<\/i><\/p>\n<\/div>\n<div>\n<p>[38] <i>See, e.g.<\/i>, <i>In re<\/i> Actos (Pioglitazone) Prods. Liab. Litig., No. 6:11-md-2299, 2012 U.S. Dist. LEXIS 187519, at *20 (W.D. La. July 27, 2012).<\/p>\n<\/div>\n<div>\n<p>[39] Global Aero. Inc. v. Landow Aviation, No. CL 61040, 2012 Va. Cir. LEXIS 50, at *2 (Apr. 23, 2012).<\/p>\n<p>[40] Kleen Products, LLC v. Packaging Corp., No. 10 C 5711, 2012 U.S. Dist. LEXIS 139632, at *61-63 (N.D. Ill. Sept. 28, 2012).<\/p>\n<\/div>\n<div>\n<p>[41]<i> <\/i>EORHB v. HOA Holdings, Civ. Ac. No. 7409-VCL (Del. Ch. Oct. 15, 2012), 2012 WL 4896670, <i>as amended in a subsequent order<\/i>, 2013 WL 1960621 (Del. Ch. May 6, 2013).<\/p>\n<\/div>\n<div>\n<p>[42] <i>In re <\/i>Biomet M2a Magnum Hip Implant Prods. Liab. Litg., No. 3:12-MD-2391, 2013 U.S. Dist. LEXIS 84440, at *5-6, *9-10 (N.D. Ind. Apr. 18, 2013).<\/p>\n<\/div>\n<div>\n<p>[43] <i>See, e.g.<\/i>, Nicholas Barry, Note, <i>Man Versus Machine Review: The Showdown Between Hordes of Discovery Lawyers and a Computer-Utilizing Predictive Coding Technology<\/i>, 15 Vand. J. Ent. &amp; Tech. L. 343, 344-345 (2013); Harrison M. Brown, Comment, <i>Searching for an Answer: Defensible E-Discovery Search Techniques in the Absence of Judicial Voice<\/i>, 16 Chap. L. Rev. 407, 407-409 (2013); Jacob Tingen, <i>Technologies-That-Must-Not-Be-Named: Understanding and Implementing Advanced Search Technologies in E-Discovery<\/i>, 19 Rich. J.L. &amp; Tech 2, \u00b6 63 (2012),<i> <\/i>http:\/\/jolt.richmond.edu\/v19i1\/article2.pdf.<\/p>\n<\/div>\n<div>\n<p>[44] <i>See <\/i>Pace &amp; Zakara, <i>supra <\/i>note 14, at 61-65.<\/p>\n<\/div>\n<div>\n<p>[45] <i>Cf.<\/i> Losey, <i>supra <\/i>note 13, at 68.<\/p>\n<\/div>\n<div>\n<p>[46] Pagan Kennedy, <i>William Gibson\u2019s Future is Now, <\/i>N.Y. Times (Jan. 13, 2012), www.nytimes.com\/2012\/01\/15\/books\/review\/distrust-that-particular-flavor-by-william-gibson-book-review.html?pagewanted=all&amp;_r=0.<\/p>\n<\/div>\n<div>\n<p>[47] Chris Snijders, Uwe Matzat, &amp; Ulf-Dietrich Reips, <i>\u201cBig Data\u201d: Big Gaps of Knowledge in the Field of Internet Science<\/i>, 7 Int\u2019l J. Internet Sci. 1 (2012), http:\/\/www.ijis.net\/ijis7_1\/ijis7_1_editorial.pdf.<\/p>\n<\/div>\n<div>\n<p>[48] Daniel Burrus, <i>25 Game Changing Trends That Will Create Disruption &amp; Opportunity (Part I)<\/i>, Daniel burrus, http:\/\/www.burrus.com\/2013\/12\/game-changing-it-trends-a-five-year-outlook-part-i\/ (last visited Feb. 24, 2014).<\/p>\n<\/div>\n<div>\n<p>[49] Bill Franks, Taming the Big Data Tidal Wave: Finding Opportunities in Huge Data Streams with Advanced Analytics 5 (John Wiley &amp; Sons, Inc. ed., 2012) (citing Stephen Prentice, CEO Advisory: \u2018Big Data\u2019 Equals Big Opportunity (2011)).<\/p>\n<\/div>\n<div>\n<p>[50] <i>Id. <\/i>at 5.<\/p>\n<\/div>\n<div>\n<p>[51] Kenneth Neil Cukier &amp; Viktor Mayer-Schoenberger, <i>The Rise of Big Data: How It\u2019s Changing the Way We Think About the World<\/i>, Council on Foreign Relations (Apr. 3, 2013), http:\/\/www.foreignaffairs.com\/articles\/139104\/kenneth-neil-cukier-and-viktor-mayer-schoenberger\/the-rise-of-big-data.<\/p>\n<\/div>\n<div>\n<p>[52] Davenport &amp; Kim, <i>supra<\/i> note 5.<\/p>\n<p>[53] <i>Id.<\/i> at 3.<\/p>\n<\/div>\n<div>\n<p>[54] <i>See<\/i> <i>id. <\/i>at 4-5 (providing a listing of various fields of research that make up a part of and comfortably fit within the broader term \u201cAnalytics,\u201d including statistics, forecasting, data mining, text mining, optimization and experimental design).<\/p>\n<\/div>\n<div>\n<p>[55] For additional titles in the popular literature, see Thomas H. Davenport &amp; Jeanne G. Harris, Competing on Analytics: The New Science of Winning (2007); Franks, <i>supra<\/i> note 49; Thornton May, The New Know: Innovation Powered by Analytics (John Wiley &amp; Sons, Inc. ed., 2009); Michael Minelli, Michele Chambers &amp; Ambiga Dhiraj, Big Data Analytics: Emerging Business Intelligence and Analytic Trends for Today\u2019s Businesses (John Wiley &amp; Sons, Inc. ed., 2013); Eric Siegel, Predictive Analytics: The Power to Predict Who Will Click, Buy, Lie, or Die (John Wiley &amp; Sons, Inc. ed., 2013).<\/p>\n<\/div>\n<div>\n<p>[56] <i>See<\/i> Davenport &amp; Kim, <i>supra<\/i> note 5<i>.<\/i><\/p>\n<p>[57] <i>See<\/i> AIIM, Big Data and Content Analytics: measuring the ROI 9 (2013), <i>available at<\/i> http:\/\/www.aiim.org\/Research-and-Publications\/Research\/Industry-Watch\/Big-Data-2013.\u00a0 In a questionnaire asking \u201cWhat type of analysis would<br \/>\nyou like to do\/already do on unstructured\/semi-structured data?\u201d, respondents identified over a dozen uses for analytics which they would consider of high value to their corporation, including: Metadata creation; Content deletion\/retention\/duplication; Trends\/pattern analysis; Compliance breach, illegality; Fraud detection\/prevention; Security re-classification\/PII (personally identifiable information) detection; Predictive analysis\/modeling; Data visualization; Cross relation with demographics; Incident prediction; Geo-correlation; Brand conformance; Sentiment analysis; Image\/video recognition; and Diagnostic\/medical.\u00a0 <i>Id<\/i>.<\/p>\n<\/div>\n<div>\n<p>[58] The Sedona Conference, The Sedona Conference Commentary on Information Governance 2 (2013), <i>available at<\/i> https:\/\/thesedonaconference.org\/publication [hereinafter <i>Sedona IG Commentary<\/i>].<\/p>\n<\/div>\n<div>\n<p>[59] Charles R. Ragan, <i>Information Governance: It\u2019s a Duty and It\u2019s Smart Business, <\/i>19 Rich. J.L. &amp; Tech. 12, \u00b6 32 (2013),<i> <\/i>http:\/\/jolt.richmond.edu\/v19i4\/article12.pdf (internal quotation marks omitted) (quoting Barclay T. Blair, <i>Why Information Governance, in <\/i>Information Governance Executive Briefing Book, 7 (2011),<i> available at <\/i>http:\/\/mimage.opentext.com\/alt_content\/binary\/pdf\/Information-Governance-Executive-Brief-Book-OpenText.pdf).\u00a0 For additional useful definitions of what constitutes information governance, see <i>The Generally Accepted Recordkeeping Principles<\/i>, ARMA Int\u2019l, http:\/\/www.arma.org\/r2\/generally-accepted-br-recordkeeping-principles (last visited Feb. 24, 2014) (setting out eight principles of IG, under the headings Accountability, Integrity, Protection, Compliance, Availability, Retention, Disposition and Transparency); Debra Logan, <i>What is Information Governance? And Why is it So Hard?<\/i>, Gartner (Jan. 11, 2010), http:\/\/blogs.gartner.com\/debra_logan\/2010\/01\/11\/what-is-information-governance-and-why-is-it-so-hard\/ (defining IG on behalf of Gartner to be \u201cthe specification of decision rights and an accountability framework to encourage desirable behavior in the valuation, creation, storage, use, archival and deletion of information. It includes the processes, roles, standards and metrics that ensure the effective and efficient use of information in enabling an organization to achieve its goals.\u201d).<\/p>\n<\/div>\n<div>\n<p>[60] <i>Sedona IG Commentary<\/i>, <i>supra<\/i> note 58, at 5.<\/p>\n<\/div>\n<div>\n<p>[61] <i>Id.<\/i><\/p>\n<\/div>\n<div>\n<p>[62] <i>Id.<\/i><\/p>\n<\/div>\n<div>\n<p>[63] <i>Id.<\/i><\/p>\n<\/div>\n<div>\n<p>[64] <i>Id.<\/i> at 25.<\/p>\n<\/div>\n<div>\n<p>[65] <i>Sedona IG Commentary<\/i>, <i>supra<\/i> note 58, at 25.<\/p>\n<\/div>\n<div>\n<p>[66] <i>Id.<\/i> at 27.<\/p>\n<\/div>\n<div>\n<p>[67] <i>See<\/i> Pension Comm. of Univ. of Montreal Pension Plan v. Banc of Am. Sec., LLC, 685 F. Supp. 2d 456, 461 (S.D.N.Y. 2010).\u00a0 The information task in e-Discovery is therefore very unlike the user experience with the leading, well-known commercial search engines on the Web in, for example, finding a place for dinner in a strange city.\u00a0 For the latter project, few individuals religiously scour hundreds of pages of listings even if thousands of \u201chits\u201d are obtained in response to a select set of keywords; instead they browse only from the first few pages of listings.\u00a0 Yet the lawyer is tasked with making reasonable efforts to credibly retrieve \u201cthe long tail\u201d represented by \u201cany and all\u201d documents in response to document requests so phrased under Federal Rule of Civil Procedure 34.<\/p>\n<\/div>\n<div>\n<p>[68] <i>See, e.g.<\/i>, Da Silva Moore v. Publicis Groupe, 287 F.R.D. 182, 191 (S.D.N.Y. 2012), <i>aff\u2019d<\/i><i>sub nom<\/i>. Moore v. Publicis Groupe SA,<i> <\/i>2012 U.S. Dist LEXIS 58742 (Apr. 26, 2012) (Carter, J.); <i>Pension Comm<\/i>., 685 F. Supp. 2d at 461.<\/p>\n<\/div>\n<div>\n<p>[69] <i>See <\/i>Bennett B. Borden et al., <i>Why Document Review Is Broken<\/i>, EDIG: E-Discovery and Information Governance, May 2011, at 1, <i>available at <\/i>http:\/\/www.umiacs.umd.edu\/~oard\/desi4\/papers\/borden.pdf.<\/p>\n<\/div>\n<div>\n<p>[70] <i>An Insider\u2019s Look at Reducing ESI Volumes Before E-Discovery Collection<\/i>, EXTERRO, http:\/\/www.exterro.com\/ondemand_webcast\/an-insiders-look-at-reducing-esi-volumes-before-e-discovery-collection\/ (last visited Feb. 24, 2014); Andrew Bartholomew, <i>An Insider\u2019s Perspective on Intelligent E-Discovery<\/i>, E-Discovery Beat (Sept. 11, 2013), http:\/\/www.exterro.com\/e-discovery-beat\/2013\/09\/11\/an-insiders-perspective-on-intelligent-e-discovery\/.<\/p>\n<\/div>\n<div>\n<p>[71] <i>See<\/i> Fed. R. Civ. P. 1 (\u201cThese rules . . . should be construed and administered to secure the <i>just, speedy, and inexpensive<\/i> determination of every action and proceeding.\u201d) (emphasis added).<\/p>\n<\/div>\n<div>\n<p>[72] Borden et al., <i>supra <\/i>note 69, at 3.<\/p>\n<\/div>\n<div>\n<p>[73] All of the \u201cTrue Life Examples\u201d referred to in this article are \u201cripped from\u201d the pages of the author\u2019s legal experience, without embellishment.<\/p>\n<\/div>\n<div>\n<p>[74] A qui tam suit is a lawsuit brought by a \u201cprivate citizen (popularly called a \u2018whistle blower\u2019) against a person or company who is believed to have violated the law in the performance of a contract with the government or in violation of a government regulation, when there is a statute which provides for a penalty for such violations.\u201d\u00a0 <i>Qui Tam Action<\/i>, The Free Dictionary, http:\/\/legal-dictionary.thefreedictionary.com\/qui+tam+action (last visited Feb. 24, 2014); <i>see also<\/i> United States <i>ex rel.<\/i> Eisenstein v. City of New York, 556 U.S. 928, 932 (2009) (defining a qui tam action as a lawsuit brought by a private party alleging fraud on behalf of the government) (internal citations omitted).<\/p>\n<\/div>\n<div>\n<p>[75] <i>See<\/i> Davenport &amp; Kim, <i>supra<\/i> note 5, at 3.<\/p>\n<\/div>\n<div>\n<p>[76] <i>See<\/i> Amended Complaint and Demand for Jury Trial, NAACP v. Balt. City Police Dep\u2019t, No. 06-1863 (D. Md. Dec. 18, 2007), <i>available at<\/i> http:\/\/www.aclu-md.org\/uploaded_files\/0000\/0205\/amended_complaint.pdf.<\/p>\n<p>[77] <i>See<\/i> Charles F. Wellford, Justice Assessment and Evaluation Services, First Status Report for the Audit of the Stipulation of Settlement Between the Maryland State Conference of NAACP Branches, et. al. and the Baltimore City Police Department, et. al. 2 (2012), <i>available at<\/i> http:\/\/www.aclu-md.org\/uploaded_files\/0000\/0207\/first_audit_report_april_30.pdf; <i>see also<\/i><i>Plaintiffs Win Justice in Illegal Arrests Lawsuit Settlement with the Baltimore City Police Department<\/i>, ACLU (June 23, 2010), https:\/\/www.aclu.org\/racial-justice\/plaintiffs-win-justice-illegal-arrests-lawsuit-settlement-baltimore-city-police-depar.<\/p>\n<\/div>\n<div>\n<p>[78] <i>See<\/i> Wellford, <i>supra <\/i>note 77, at 2, 14.<\/p>\n<\/div>\n<div>\n<p>[79] <i>See<\/i> Jason R. Baron &amp; Simon J. Attfield, <i>Where Light in Darkness Lies: Preservation, Access and Sensemaking Strategies for the Modern Digital Archive<\/i>, <i>in<\/i> The Memory of the World in the Digital Age Conference: Digitalization and Preservation 580-595 (2012), http:\/\/www.ciscra.org\/docs\/UNESCO_MOW2012_Proceedings_FINAL_ENG_Compressed.pdf.<\/p>\n<p>[80] <i>See id. <\/i>at 587.<\/p>\n<\/div>\n<div>\n<p>[81] <i>See id.<\/i> at 588; <i>see also<\/i> Ragan, <i>supra<\/i> note 59, at \u00b6 6.<\/p>\n<p>[82] <i>See, e.g.<\/i>, The Sedona Conference, The Sedona Conference Commentary on Inactive Information Sources 2, 5 (2009), <i>available at <\/i>https:\/\/thesedonaconference.org\/publication\/The%20Sedona%20Conference\u00ae%20Commentary%20on%20Inactive%20Information%20Sources.<\/p>\n<\/div>\n<div>\n<p>[83] <i>See<\/i><i>Sedona IG Commentary<\/i>, <i>supra<\/i> note 58, at 20-22.<\/p>\n<\/div>\n<div>\n<p>[84] <i>See<\/i> Davenport &amp; Harris, <i>supra<\/i> 55, at 176-78.<\/p>\n<\/div>\n<div>\n<p>[85] The Sedona Conference, <i>The Sedona Conference Commentary on Achieving Quality in the E-Discovery Process<\/i>, 10 Sedona Conf. J. 299, 325 (2009).<\/p>\n<\/div>\n<\/div>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>DownloadPDF Cite as: Bennett B. Borden &amp; Jason R. Baron,\u00a0Finding the Signal in the Noise: Information Governance, Analytics, and the Future of Legal Practice, 20 Rich. J.L. &amp; Tech. 7 (2014), http:\/\/jolt.richmond.edu\/v20i2\/article7.pdf. \u00a0 Bennett B. Borden* and Jason R. Baron** \u00a0 Introduction [1]\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 In the watershed year of 2012, the world of law witnessed the [&hellip;]<\/p>\n","protected":false},"author":4287,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"jetpack_post_was_ever_published":false,"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":false,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2}},"categories":[1228],"tags":[],"class_list":["post-2017","post","type-post","status-publish","format-standard","hentry","category-articles"],"jetpack_publicize_connections":[],"jetpack_featured_media_url":"","jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/paMHOZ-wx","jetpack-related-posts":[],"_links":{"self":[{"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/posts\/2017","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/users\/4287"}],"replies":[{"embeddable":true,"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/comments?post=2017"}],"version-history":[{"count":0,"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/posts\/2017\/revisions"}],"wp:attachment":[{"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/media?parent=2017"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/categories?post=2017"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blog.richmond.edu\/jolt\/wp-json\/wp\/v2\/tags?post=2017"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}