Research data lost to the sands of time

[Originally posted on Free Government Information blog] Here's an interesting article, not on link rot (a topic FGI has been tracking for some time), but on *data rot*. In a recent article in Current Biology, researchers examined the availability of data from 516 studies between 2 and 22 years old. They found the following:

  • that the odds of a data set being reported as extant fell by 17% per year;
  • Broken e-mails and obsolete storage devices were the main obstacles to data sharing
  • Policies mandating data archiving at publication are clearly needed

Librarians have known of this issue for years -- the Inter-university Consortium for Political and Social Research (ICPSR) was set up in 1962 to tackle this -- but it does put the issue in focus. And finally the federal government -- via efforts like the NSF’s data management plan and OSTP's new directive to improve the management of and access to scientific collections -- is beginning to get behind the effort to improve on data rot. And many libraries — not to mention scientists and researchers — are beginning to struggle with the issue of data preservation. The issue is too big for just government information librarians to handle obviously. But this is fertile space in which govt information librarians, data librarians, research communities, and federal agencies can come together. The Federal policy stating the importance of data preservation is there, it’ll just take effort by multiple stakeholders to make sure it actually happens. It's a positive that the writers of Dragonfly, the blog of the National Network of Libraries of Medicine Pacific Northwest Region -- where I came across the article -- point out that academic institutions can and should play a leading role in data preservation. I wholeheartedly agree! Vines, Timothy H., Arianne YK Albert, Rose L. Andrew, Florence Débarre, Dan G. Bock, Michelle T. Franklin, Kimberly J. Gilbert, Jean-Sébastien Moore, Sébastien Renaut, and Diana J. Rennison. "The availability of research data declines rapidly with article age." Current Biology 24, no. 1 (2014): 94-97.

