top of page

The following is a guest post by Jefferson Bailey, Program Manager and Partner Specialist at Internet Archive and co-chair of the Innovation Working Group of the National Digital Stewardship Alliance.

At July’s Digital Preservation 2014 conference, hosted by the Library of Congress in Washington D.C., a team from the National Digital Stewardship Alliance (NDSA) Innovation Working Group held a session to formally launch the Digital Preservation Questions and Answers online forum. As the NDSA starts to engage with other associations to help spread the word about this resource, and having recently presented at the LIPA business meeting at AALL, we were glad when Margie invited us to write up a blog post for the LIPA site describing the Q&A project and website. LIPA is an active NDSA member and strong proponent of digital preservation and this is the perfect place to begin reaching out to NDSA members and affiliated organizations to promote this project and encourage use and participation. In this post, I’ll give a quick “who, what, when, where, and how” about the Digital Preservation Q&A site. For background on how the project originally came about, see this prior blog post: http://blogs.loc.gov/digitalpreservation/2014/07/digital-preservation-questions-meet-digital-preservation-answers/

What is the Digital Preservation Q&A site?

The Q&A site aims to provide a single, user-friendly forum for exchanging information, professional practices, and individual and institutional knowledge around all matters related to digital preservation. As a field that evolves alongside technological and social change, digital preservation is especially reliant upon professionals sharing their wisdom and practices with others; however, this information is often exchanged on listservs, Google Groups, social media, and other informal networks and channels that can be hard to find, search, or share, or preserve. The Digital Preservation Q&A forum was created to be a unified, public, searchable, place to share this information on a platform that is hosted and administered by groups dedicated to supporting digital preservation practice, education, and knowledge sharing.

Who participates in the site?

The forum currently is being used by a wide diversity of those interested in digital preservation. Credentials are not needed to participate! Already the site reflects a range of institution types, levels of expertise, and topical areas. As the site’s about page states, the forum is generally intended for “information technologists, archivists, engineers, librarians, computer scientists, curators, web developers and others to help each other make best use of tools, techniques, processes, workflows, practices and approaches to insuring long term access to digital information.”

What sort of Q&As are relevant?

Questions and answers can range from general advice seeking to highly technical recommendations — the site intends to be an open forum, friendly to the information-seeking and advice-giving general public as well as the experienced practitioner. Questions have covered topics as diverse as policy development, cost models and hardware prices, migration and emulation, file formats, web archiving, and digital forensics.  While the site does aim to mediate questions that strongly associate digitization with digital preservation, it is open to all question and answers that involve digital preservation. Questions about questions are welcome too and can be found via the “meta” tag. In the planning stages, the forum’s creators explicitly decided to take a free-wheeling, figure-it-out-as-we-go approach towards allowing questions, so interested users are encouraged not to feel daunted in asking their digital preservation questions. All questions and answers are welcomed!

Who administers the sites?

The site was started by a collaboration between the NDSA, the Open Planets Foundation (OPF), and interested practitioners. It is jointly managed by NDSA & OPF and hosted on servers maintained by the OPF. To mitigate against potential spam, creating a user account is required, but remember that the site is hosted on the infrastructure of OPF, an organization “established to provide practical solutions and expertise in digital preservation.” The site will not be bound to the whims of corporate profit-seeking.

How you can help

Asking and answering questions, or even commenting on specific questions and answers, is an obvious, easy way to contribute to helping build the site. But you can also contribute just by using the site as an informational resource. The “voting” feature allows to you recognize question and answers you find particularly informative or insightful. Though the forum is new and the vote counts currently appear to be low, as the pool of information grows, and use of the forum expands, the value of user voting will become more and more apparent as it helps filter the most relevant questions and answers to the top. Be sure to bookmark and visit the Digital Preservation Questions and Answers site and start helping to build this vital digital preservation community resource!

 
 
 

The following is a guest post by Nicholas Taylor, Web Archiving Service Manager for Stanford University Libraries and former Library Technology Specialist for the U.S. Supreme Court.

The topic of link rot in legal materials has been an area of study for some time (PDF) but has only recently surfaced in the mainstream press with the high-profile findings by two separate law journal articles of the high incidence of broken links in U.S. Supreme Court opinions. It’s a problem that the Court has made some effort to mitigate, albeit in a disjointed way. Staff working for the Reporter of Decisions pre-check the citations in all-but-published opinions and save cited web content to PDF for inclusion in the files, an approach consistent with 2009 guidance (PDF) issued by the U.S. Federal Courts. The Ninth Circuit Library uses a similar technique but does one better by actually posting the PDFs online.

The current setup isn’t ideal, though. A cited link in a PDF opinion provides a false sense of confidence that the resource it once represented still exists at the same location and, more critically, still exists at the same location unchanged. What might it take to put in place a solution for more trustworthy archiving?

Let’s unpack the problem a bit more.

PDFs of web cites (adopting the Ninth Circuit Library’s handy neologism) are created late in the process, so there’s no guarantee that the represented resource is the same as what the opinion drafter intended, potentially many months prior. This was cheekily pointed out by the new owner of the host http://ssnat.com/ after Justice Alito cited it in his concurring opinion for Brown v. Entertainment Merchants Association. Though presumably the Reporter’s Office stores a PDF that more accurately reflects that particular web cite, the Court can make no general assurances that the stale web cites within the opinions themselves are still accurate or even functional.

Archiving links at the same time they’re cited would optimally allow for the web addresses of archived versions to be placed along side the original web cites in the final opinions, but it presents challenges. The archiving process would need to be simple, so as not to burden opinion drafters. While creating a PDF from a web page is trivial, creating a workflow to aggregate and manage the many resulting PDFs would be hard. Protecting the confidentiality of the Court’s research imposes additional requirements. The archiving requests of the different chambers would need to be firewalled from each other and the rest of the Court. If WARC rather than PDF were the format of choice, automated methods of archiving web cites like Heritrix or wget would need to be carefully disguised or proxied through a third-party service provider, so as not to leak hints about which way the Court may be leaning or along what lines of argument.

Strong objections would need to be overcome to employ a third-party service provider for a purpose so close to the Court’s research. On the other hand, recent transparency-focused applications of government agency IP address ranges (e.g., @congressedits) might make outsourcing more appealing. Reed Archives would be a logical candidate, as they are a subsidiary of LexisNexis, a company that the Court already trusts. Perma.cc would be another candidate. The multi-institutional storage redundancy is likely to be seen as more of a liability than a feature in this context, though, and provisions would need to be established to ensure that copyright claims didn’t trump the Court’s prevailing interest in providing persistent access to the bases for its jurisprudence.

With its browser integration, click-to-capture ease-of-use, and focus on citation metadata management, Zotero would both readily handle the data capture requirements and integrate with the research workflow. Data captured to individual Zotero libraries could be aggregated by means of WebDAV syncing to an internal server. Unfortunately, WebDAV syncing isn’t supported for Zotero Group libraries, which would likely be essential for the Reporter’s staff to corroborate synced files with their originating web cites and might otherwise be used to good effect for flexible, discrete, and discreet collaborations within and/or beyond a given justice’s chambers. Even if this were a feature, all groups must be registered and managed through the Center for History and New Media, posing additional overhead and possible confidentiality issues.

If a solution for near-contemporaneous web archiving could be configured, then the resulting archived web cites could be made accessible via permalinks pointing either to PDFs or, in the case of WARCs, Wayback access points hosted by the Court, and embargoed for public access until the opinions were published. Foreknowledge of the “permanent” location of the archived web resources would allow the Reporter’s Office to amend the opinion web cites with the permalinks prior to publication. To be sure, there are many more solutions available if archiving took place after the opinions were published, but these web cites couldn’t be treated as authoritative. With consideration of some of the challenges outlined above, I hope the Court can progress toward ensuring that their opinion web cites are as reliable as anything else referenced in their opinions, without compromising the confidentiality of the research process.

 
 
 

The Utah State Library, a division of the Utah Department of Heritage and Arts, has awarded the S.J. Quinney Law Library at the University of Utah,  an LSTA (Library Services and Technology Act) grant of $45,181 to digitize Utah Supreme Court briefs from 1929-2000 and make them publicly accessible on the web.

Through the Utah Supreme Court Briefs Digitization Project, the Quinney Law Library, along with its partner, the Howard Hunter Law Library at Brigham Young University, will expand on the development of the Hunter Law Library’s existing Utah Court Briefs Repository, by incorporating another 6,500 briefs to the collection. The Hunter Law Library has generously agreed to host these additional materials in its Bepress repository, which provides an easy to use web interface to assist lawyers, legal educators, and legal historians.

An added benefit to the project is statewide visibility to the briefs through the Mountain West Digital Library (MWDL) digital collections portal, which harvests BYU’s Utah Court Briefs Repository. Through MWDL, the collection gains national recognition, by the briefs being visible in the Digital Public Library of America (DPLA), which harvests MWDL.

Anyone with questions about the project is welcome to contact Valeri Craigle at the Quinney Law Library, who is heading up the project.

 
 
 
bottom of page