In spring 2024 an expiring travel voucher provided me with an unusual opportunity in the post-Covid era — I was able to visit colleagues in person.1 Since my research focuses on historical data, I made a trip to the University of Minnesota for a few days to visit the librarians and archivists behind several publications on analog data curation and reuse. We had never met in person, but we had collaborated as members of the Data Curation Network (DCN) and on a webinar about analog data that led to an article.2 Even though we already knew each other, and I was already steeped in the literature on analog data, my in-person visit yielded valuable insights into the curation process and surprised me with unpredictable benefits that can only occur during an extended stay.
When most people think of data curation, they think of digital data, and most of the time they would be right. The DCN is an expertise hub for data curators working with myriad forms of digital data at web-based repositories across the US, and it is primarily a virtual operation (Johnston et al. 2018). Although we have an annual in-person meeting, members communicate year-round through virtual meetings, email, and Slack. In fact, a core function of the DCN, our curation exchange, can only work virtually. Repositories typically employ a small number of curators but accept a wide variety of data, which means most curation teams have knowledge gaps. When a DCN member institution receives a data deposit that falls outside their expertise, they can send up a virtual Bat-Signal to the entire network for support. Whoever has the necessary time and skills can help by providing guidance, tech support, or curation services. In the DCN we refer to this skill sharing as part of our radical interdependence (Carlson et al. 2023). With a diverse network of distributed teams, using all the virtual tools at our disposal, we can cover those knowledge gaps and provide robust curation support for data from a wide range of disciplines.
Most data curation is digital, but data collection started well before the digital age, and there are a few of us in the DCN who have one foot in the analog world. For the purposes of this discussion, by analog data I mean data recorded on paper. These records could be handwritten, typed, printed, or some combination thereof. Historical data, also sometimes referred to as heritage or legacy data, is an overlapping category that also includes digital data at risk of loss due to its age. Analog data can provide rare glimpses into the past that propel science forward — think weather records from old ship’s logs that provide data for climate models, or dusty field notebooks that show how decades of agriculture slowly impacted soil fertility — but this kind of data presents complex challenges for curation and reuse.
The first major hurdle is locating analog data. Searching for data is not typically part of the digital data curation process. Digital data producers who deposit their data are motivated by funder mandates, journal policies, or open science principles. Those of us who curate analog data usually must go looking for it — and it is not an easy search. Investigations at Minnesota revealed that analog data records are often scattered across campus, that usually only a few people even know the data exist, and when data do make it to the University Archives they are rarely labeled as data, which significantly complicates discovery (Farrell et al. 2023; Farrell et al. 2020; Farrell and Kelly 2018). This is consistent with my own experience as part of a data curation working group at the University of Illinois Urbana-Champaign. Our multi-year search for agricultural data from the historic Morrow Plots experiment required extensive digging in the library and archives as well as interviews with current and former project stewards (Anderson et al. 2023; 2022). Once paper data records are located, the analog data curator faces thorny questions of interpretation, preservation, and access — questions I gained insight into during my trip.
Luckily, my visit to Minnesota coincided with a scheduled appointment to collect analog data from the Horticultural Research Center, the birthplace of several widely known apple varieties including the Honeycrisp (University of Minnesota’s College of Food, Agricultural and Natural Resource Sciences 2024). Librarians and archivists have worked with the Center since 2016 to curate, document, and digitize fruit breeding records dating back to the Center’s founding in 1908 (Farrell et al. 2019). The project centers on “the vault,” a closet-sized room guarded by an old-fashioned vault door and filled with bound notebooks, half-page binders, brad-bound books of notecards, file folders, boxes, just about any form of paper record imaginable (Figure 1 and 2). We were there to pick up the last batch of materials to be digitized — a few dozen half-page binders and brad-bound books holding data mostly about apples but also strawberries, currants, and other fruits. One of the lead researchers at the Center had already pulled the materials for us, so we went ahead with the next step of the process — counting the number of pages in each resource and noting if they were single-sided, double-sided, or both — to assist the digitization lab that would be scanning them.

Figure 1 and 2: The “vault” door and a sample of the analog data records contained inside.
Paging through these paper records provided a fascinating glimpse into how the data had been collected decades ago. Penciled dates, notes, hash marks, and sketches covered custom forms printed on heavy stock firm enough to be carried out into the fields. I could see where planning did not quite line up with reality as some sections of the forms were consistently left blank and others were repurposed or supplemented with notes in the margins. It is easy to think of data collection as rote and technical, especially when forms are involved, but as I paged through these books, I saw evidence of both the wildness of plant life and the creative choices of the scientists documenting it. There is art in these scientific records, both in the decisions to break with the prescribed format when needed and more literally in the delightful illustrations such as those in the 1897 Apples book digitized by the University of Minnesota Libraries (Green 1897). Idiosyncrasies like these make analog data charming and fascinating, but they also make it difficult to translate the data into a shareable, reusable format (Figure 3).

Figure 3: An example of sketches of apples drawn by University of Minnesota researchers.
We had a chance to talk through some of these difficulties after we finished our page counts. All of us have worked with a wide variety of data, and we agreed that interpreting analog data really does take more time — more time than interpreting contemporary digital data records and more time than we tend to expect because it is hard to know what mysteries are lurking in old paper records. It has only been in the past couple of decades that most researchers have started to view data as publishable. The analog data records we have curated tend to be more like private, rough drafts filled with handwriting, symbols, and shorthand that can be hard to decode. Analog data records usually do not include the kind of explanatory documentation that makes published data understandable and reusable. We agreed that, in the absence of documentation, we needed to spend significant amounts of time sitting with the data, getting to know it well enough to begin to see the patterns and meaning that documentation might have provided. Then, writing new documentation from scratch sometimes requires extensive research in libraries and archives, which adds even more time to the curation process.
Lack of documentation is challenging enough on its own, but then there’s also the disorder. Analog data does get reused, but it tends to get passed down through generations of researchers, and that is a messy process. By the time I visited the Horticultural Research Center, the analog data records were neatly arranged and inventoried, but it was in disarray when the Minnesota team first found it. It took a team of curators three full days to organize and inventory the closet-sized collection (Farrell et al. 2019). Useful analog data records do not sit in their original order gathering dust and waiting for curators to discover them. They get handled over and over again. The more useful the data, the more it gets shared and reshuffled, relabeled, repurposed, and divorced from its original context, complicating our work as curators.
Actually, disorganized analog data is much easier to work with than data that is missing altogether. Sometimes the more useful the data, the harder it is to locate. At the Center, even before we started counting pages, we could tell that some records were missing from the batch left to be digitized. These data records had been borrowed by a researcher, which is fantastic! Curators spend so much time working with data hoping that someone will put it to good use. However, this experience made me wonder if the most useful analog data might be the last to be curated because researchers cannot afford to part with it. It’s easy to submit a copy of a digital data file for curation, but if one unique notebook is useful, how long will it be before researchers are willing to hand it over to be digitized?
Working with analog data is slow. It is cryptic. It is messy, and the most useful data might be the hardest to find. It is also rewarding and magical and working alongside other curators helped clarify my thinking on the topic in ways I didn’t expect. Being away from my regular work environment, even for just a few days, allowed me to focus on analog data without worrying about my inbox or all the other daily demands of academic life. Helping colleagues with their data allowed me to see patterns that apply to more than just my own work. Unstructured time spent chatting with those colleagues helped turn those patterns into insights and new research questions. While I greatly appreciate the accessibility of virtual collaboration tools, there is something to be said for taking advantage of opportunities to work together in person.
References
Anderson, Bethany G., Erin Antognoli, Sandi L. Caldrone, Justin D. Derner, Shannon L. Farrell, Katrina Fenlon, John R. Hendrickson, et al. 2024. “Issues and Paths Forward in the Identification and Reuse of Historic Analog Records.” Frontiers in Environmental Science 12 (April). https://doi.org/10.3389/fenvs.2024.1338628.
Anderson, Bethany G., Sandi L. Caldrone, Joshua Henry, Heidi J. Imker, Hoa Luong, Kelli Trei, and Sarah C. Williams. 2022. “Cultivating the Scientific Data of the Morrow Plots: Visualization and Data Curation for a Long-Term Agricultural Experiment.” In Proceedings iPRES 2022, 12-16 September 2022, 202–7. Glasgow, Scotland: Digital Preservation Coalition. https://doi.org/10.7207/ipres2022-proceedings.
Anderson, Bethany G., Sandi L. Caldrone, Joshua Henry, Heidi J. Imker, Andrew J. Margenot, and Sarah C. Williams. 2023. “Publishing Agricultural Data from the Morrow Plots: The Value and Logistics of Preserving a Long-Term Research Experiment.” In 19th International Conference on Digital Preservation (iPRES) Proceedings, 19-23 September 2023, 127–36. Urbana-Champaign, Illinois: University of Illinois. https://hdl.handle.net/2142/121102.
Carlson, Jake, Mikala Narlock, Mara Blake, Joel Herndon, Heidi Imker, Lisa Johnston, Wendy Kozlowski, et al. 2023. The Art, Science, and Magic of the Data Curation Network: A Retrospective on Cross-Institutional Collaboration. Michigan Publishing Services. https://doi.org/10.3998/mpub.12782791.
Farrell, Shannon, Julia Kelly, Lois Hendrickson, and Kristen Mastel. 2023. “A Pilot Study to Locate Historic Scientific Data in a University Archive.” Issues in Science and Technology Librarianship 103 (May). https://doi.org/10.29173/istl2728.
Farrell, Shannon L., Lois G. Hendrickson, Kristen L. Mastel, and Julia A. Kelly. 2020. “Historical Scientific Analog Data: Life Sciences Faculty’s Perspectives on Management, Reuse and Preservation.” Data Science Journal 19 (1): 51. https://doi.org/10.5334/dsj-2020-051.
Farrell, Shannon, Lois Hendrickson, Kristen Mastel, Katherine Allen, and Julia Kelly. 2019. “Resurfacing Historical Scientific Data: A Case Study Involving Fruit Breeding Data.” Journal of eScience Librarianship 8 (2): 1171. https://doi.org/10.7191/jeslib.2019.1171.
Farrell, Shannon L., and Julia Ann Kelly. 2018. “Identifying Potential Solutions to Increase Discoverability and Reuse of Analog Datasets in Various Campus Locations.” Issues in Science and Technology Librarianship 88 (March). https://doi.org/10.29173/istl1717.
Green, Samuel Bowdlear. 1897. Apples. https://umedia.lib.umn.edu/item/p16022coll320:475.
Johnston, Lisa R., Jake Carlson, Cynthia Hudson-Vitale, et al. 2018. “Data Curation Network: A Cross-Institutional Staffing Model for Curating Research Data.” International Journal of Digital Curation 13 (1): 125–40. https://doi.org/10.2218/ijdc.v13i1.616.
University of Minnesota’s College of Food, Agricultural and Natural Resource Sciences. 2024. “Horticultural Research Center.” Minnesota Landscape Arboretum. 2024. https://arb.umn.edu/HRC.
Thanks to the Data Curation Network for generously supporting the trip and for being flexible when the weather interfered with my original plan to attend their annual meeting, leaving me with a time-limited flight voucher.↩︎
Anderson, Bethany G., Erin Antognoli, Sandi L. Caldrone, Justin D. Derner, Shannon L. Farrell, Katrina Fenlon, John R. Hendrickson, et al. 2024. “Issues and Paths Forward in the Identification and Reuse of Historic Analog Records.” Frontiers in Environmental Science 12 (April). https://doi.org/10.3389/fenvs.2024.1338628.↩︎