(RESPONSES DUE Friday, February 26) EaaSI Hosted Pilot Forum Discussion

Greetings @us-hosted Cohort Participants!

Thank you all for your patience and graciousness over the last week or so as we dealt with weather conditions that Texas was not prepared for!

This week’s activity is our first forum discussion.
First some food for thought:

Please respond to the discussion questions below by replying to the thread with your thoughts and experiences, any follow-up questions that you want to share, and/or responses to your colleagues’ posts.

Now for our discussion questions:
1. Given the varied granularity and modes of documenting and describing software provided in your colleagues’ responses to the stakeholder questionnaire, can you think of 1-3 approaches that we could consider as a community to better understand what already exists in our collections versus what might require a coordinated collection approach?
----> Different levels of granularity and modes for describing software:

  • Modified version of the Media Archeology Lab schema
  • Qualified Dublin Core in a collection management system or other descriptive system
    • ArchivesSpace
    • Islandora
    • DSpace
  • Multiple databases that are collection-specific
  • MARC cataloging for media and software
  • Planning to implement SMRF from SPN Metadata Working Group
  • We do not catalog most of our users’ software - but we provide guidance on the metadata that research software users and engineers should use when submitting their software to GitHub or other Open Science web platforms
  • EAD (finding aid)
  • For supplementary description, we record oral histories
  • Intake/curation interviews with researchers or research software engineers
  • Recently undertook a project to review existing collections that have only been described in the aggregate in order to identify software
  • Inadvertently collected and under described - ingested in bulk as part of transfer from government departments to the EDRMS

2. What do you all think about a developing a coordinated documentation and collection strategy among members of SPN and EaaSI nodes? What might that look like?

3. Looking at the five primary modes of access for born-digital materials among members of the cohort, have any you performed time trials for some of your access modes or otherwise tried to measure effort and risk associated with each of these approaches? What is the level of efficiency that your organization hopes to reach with EaaSI - what level of service efficiency would make the adoption of EaaSI service worthwhile (in your mind at this early stage - prior to formal testing of the platform)?
----> Five primary modes of access for born-digital materials:

  • Continuity of service + security risk - keeping legacy web applications on life support
  • On-premises legacy workstations
    Repositories and online discovery systems
    Cloud-based file sharing
  • Emulation or something like it

4. Given the group consensus on the wide range of stakeholders that are implicated in the adoption of a software preservation strategy and emulation services, how are you all thinking of engaging these groups if they are not already represented as part of your pilot Node team?

Hi @us-hosted cohort participants!

Hope you are having a great week so far!

Just a friendly reminder to offer up any thoughts or follow-up questions for our first forum discussion!
The first question, on understanding what is already in our collections in the service of reducing redundancy of effort, sets us up nicely for the next activity - A spot-check software inventory - which I will post for next week.

The second question, on coordinated software collection development for certain types of software, gets us thinking about the role we want our community to play in bridging some scale/scope challenges we face in terms of dependencies.

The third question, on involving stakeholders beyond the members of your current Node team, poses an opportunity to help learn from each other’s approaches to advocacy and securing organizational investment in this work.

Feel free to respond to one or all of the questions! Remember, we are pushing much of our discussion to asynchronous engagement so that people can participate at a time during the week that works best for them so please do share your reflections and experiences — we’ll revisit these forum discussions after formal testing when we are setting up roundtables.

Thank you all!

Yes! Feel free to reply to the thread - thank you for asking, Cynde!

Great questions @jmeyerson ! I have some thoughts on questions 2.

I think a coordinated documentation/collection strategy would be an excellent approach to building a shared “collective collection.” It seems like a natural extension of similar work we do at our institution in league with partner institutions to help reduce replication of work, minimize collection footprints, and increase budget efficiency. And y’all have already created such an excellent cohort of contributors!

One idea offhand could be to create a registry (similar to what Cobweb attempted for web archives) where folks could express needs and assets. People could nominate software/version that was needed, as well as any manuals/documentation that would be helpful, and others could claim whatever aspects they could contribute (or plan to acquire and then contribute).

2 Likes

Question 1 and 2 - I agree with @tricia_patterson that a coordinated documentation/collection strategy would be quite useful. I also like the idea of a registry, but I wonder if it might be possible to build a link to it into EaaSI (or even build it into EaaSI!) so I can search or browse for software and manuals I might be needing as I build my own environments in my local instance. I like things like this to be at my fingertips in the system I am using if possible, rather than having to remember to go to another web location to find what I want. Almost like having an EaaSI marketplace (but not a market as it wouldn’t be for sale, but you could track popularity of items and downloads). Additionally, you could make it like Reddit, where users register all of their items that they are willing to share, without necessarily having uploaded them, and other users could then up-vote for certain popular items to be digitised (into disk images, vhd or ISO) and uploaded if they get enough demand in up-votes. Not saying we should gameify EaaSI, but it would be an interesting prospect.

Question 1 and 2 both have me asking, “How do I know what environments other institutions have without trawling through each institution’s published environments and objects list?” This will become more important as more institutions join EaaSI (which I hope will happen).

Question 3 - As yet we at the National Archives of Australia only have a local non-networked version of EaaSI from here - Try EaaSI. We haven’t yet received support internally to scale up to this version - Setup and Deployment so we haven’t performed any time trials as such, although I can tell you it does take me quite a few attempts and many days sometimes to build an environment I can work with to open certain items in our collection. Even then, it still isn’t perfect and I am finding errors like video cards not being detected causing issues with the graphics, to sound not playing and this comes down also to my lack of knowledge about how to set each of these environments up, not being a trained IT person. But what we have used it for, ie opening objects in obsolete file formats just to look at them for appraisal purposes, has worked and has been quite beneficial.

The level of efficiency I would like to get to will rely on my institution’s ICT support, but also we found the latency of using the public beta online test version from Australia, to be unusable. Even the local copy can be slow and the mouse doesn’t follow the actual mouse on the screen at all times. I would like to get it to a point where we can use an internally networked version of EaaSI (as in the deployed version), that is capable of being easily able to import items (maybe even linked straight from Preservica - which I hear could be in the pipeline), view them, manipulate them somehow if necessary and spit those manipulated versions back out into the modern Windows 10 environment if that is possible, or if not, then be able to have our general public click on an item, and to paraphrase @ecochrane “automagically” opens in a viewer online. We need to be able to conduct access examination and redaction on some of these items, which will be difficult if that redaction might have to happen at the coding level. I want it to be able to handle large emulations of networked environments with interlinked oracle databases produced by our government departments, like massive petabyte databases that call out to millions of Large Objects. I want to have an easy solution for our government departments to take full hard disc images easily of their servers so we can replicate them when we take in snap shots of their systems.

But then there is what I would like, and what we can feasibly support at the NAA and within the 180 or so Australian Government public departments and agencies. Even if we can just demonstrate internally how to emulate a mid-tier networked database like our own RecordSearch, then that might convince some people at the NAA this is a prospect worth pursuing. At present all they see it as useful for is the use case I have been able to present them with which is what we have done with it (ie open objects in obsolete file formats).

In terms of question 4, the only way we can engage our stakeholders is through proof of concept and demonstration. We need to prove to our Executive and ICT it isn’t overly difficult or expensive to implement, and that there is archival value in doing so, and not doing it means loss of Australian historical and cultural data. This might mean step by step guidelines on implementation, and possibly even some initial hands on assistance to show our stakeholders how to implement it.

I understand we are a bit of a different animal in being a national archive which is both a cultural institution and a democratic access institution with data management responsibilities, as opposed to a university or a library or a museum.

I look forward to other people’s responses.

Tim Mifsud
Digital Archivist
National Archives of Australia

1 Like

I’ve been wondering whether what existing naming / organization we might be able to use to coordinate our collections. For instance, is something like NIST’s Common Platform Enumeration helpful? (They have a search feature, as well.)

For virtual machine specification, do we draw on anything like libvirt? Would that be helpful or not?

I realize that we might end up crafting a sort of homegrown system, but perhaps we don’t have to reinvent all of that wheel?

I’ve also been thinking about how to draw in stakeholders, and my initial feelings right now is that, right now, there is no specific service model around software preservation and/or access to emulated environments. With research data, we operated in this space of providing ad-hoc support (which is where we’re at right now with software preservation I think) until a sort of triggering moment happened. For research data, this was the moment when NSF announced that they would start requiring data management plans in all grant proposals. I keep wondering, is there a similar moment for software preservation? Or has one already happened for some folks? If not, what might it be? What would that look like?

I guess – this is a good question for the group: if your organization has had that moment, what was it? What happened?

1 Like

In response to question 3, I haven’t in my experience performed time trials or otherwise measured risk when providing access to born-digital materials, nor has Houghton, due the nascent nature of our work in terms of access. However, I do think measuring risk and/or doing a usability study across different modes of access for sets of interactive objects be extremely useful.

In my mind, I imagine that we’d have some sort of programmatic access to EaaSI that would my library (Houghton) to have their own instance to be used both for public access and for staff use (i.e. appraisal). It would be lovely to have some sort of integration with local systems that we share amongst nodes, institutions, etc. OR with cloud-base file sharing systems. I’m interested what other units, particularly at Harvard, think, as well as other cohort members!

@dianne.dietrich I agree that a “trigger moment” of some sort needs to happen in order to a) incentivize investigating stakeholder interest (i.e. make it worthwhile), b) create interest, and c) ultimately move toward that common goal. Gauging interest can be such a huge lift depending on the scope and size of the institution since software preservation is so broad. At University of Arizona, we used the Fostering Communities of Practice grant project as an opportunity to interact with the broader campus and discuss engagement or need with software preservation. The goal of that was to sustain interest/engagement in a new SPN membership, while also investigating what tools, strategies, or service models could help our institutional needs (if we were to go further with EaaSI, what other stakeholders might it work for?). I think it’s a bit of a chicken and the egg situation but gauging light-weight stakeholder interest in the beginning can be helpful in planning out/right-sizing those service models, particularly in financially precarious times.

Oh, I love this registry idea! It reminds me vaguely of the EaaSI issue tracker in the Fostering Communities of Project tracker but extended. +1

Regarding the third question (involving stakeholders beyond the members of three current Node teams), we definitely view this pilot activity as proof-of-concept work that will lead to offering EaaSI services across the Library as well as appropriate units across the University, e.g., instructional computing, research technology, etc. We view software preservation as becoming a core service imperative of our digital preservation program. Without it, we cannot be effective in meeting the second of our twin goals of assurances regarding the persistence of authentic information objects and the persistence of opportunities for legitimate information experiences. The pilot projects will let us work through the pertinent issues and help us design our eventual programmatic initiative.

Throwing my opinions into the ring:

third question first: This is our second pass at using EaaSI, and our second EaaSI team. We chose the team this time around with an eye to targeting some specific library teams we could work with to promote EaaSI as a service. Our current team contains members from some of the targeted user communities we think are willing to champion software preservation at our institutions: namely archives and special collections. I think in our first pass we cast a wide net to show the viability of emulation as a modality for preservation, but we made little traction building the case that there were communities served by our library that really wanted to use this service. Now that we have communities who have stepped up and said “I need this service”, my highest priority during the pilot is to implement specific use cases for those communities that they (and I) can use to sell the idea to my library management that EaaSI should be part of our developing digital preservation policy.

Second question: Our proposed collaborative infrastructure absolutely requires that we adopt the same metadata schema to describe what we are putting into EaaSI, and I know some of the nodes are way ahead of us in figuring this out, and I hope we can learn from what they have done.

First Question: I will be honest- having spent quite a bit of time on the software inventory during the EaasI grant work, I dread any more inventory work. I see value in knowing which node partners share our interests, and being able to utilize some of their work, but I put getting my “sales pitch” use cases done first, else I may not be a partner for long.

1 Like

Hi everyone,

Here are my thoughts on the second question: An emphatic YES to developing a coordinated documentation strategy. Thinking outloud I wonder if using the community forums is a great way to at least begin to think and model how we might want to do this. My rational for advocating for a coordinated documentation is that with Stanford’s experience so far there is no way that one single institution can manage to document all of the neccessary workflows, steps, dependancies. we need to find a way to better leverage the knowledge and expertise being developed at our peer partners.

The question of how we might coordinate our collecting activities is interesting one. As more of a technical lead for our node team I have my own ideas as to what ‘complimentary’ software we need to collect to facilitate access to collections that get delivered to our lab. In other words I typically look through a collection quickly to get a sense of companion applications that we may need to be able to do something with collection materials. As for how or why we systematically go after and acquire specific collections I have no detail insight into our collection development plans and at our instiution we are not strong on publishing and talking about our collection activies before they happen. What about using EaaSI itself as an automated way of gethering this sort of data? This is likely openning up a big can of worms with this suggestion but I wonder if there isn’t a way to use the existing framework to get some transparency on the software our peers are trying to preserve.

Question 3. We have not done any timings to determine the level of work for setting up emulated environments but would love and need to do this if we are going to advocate for using EaaSI in a production service. Our use cases have been too different across the collections we’ve been working with to really get a sense of timings which = costs. We’re closest to doing this with our IMF collection but I’d love to know how others are doing this.

Programatically I can’t see how Stanford can serve its research community without providing some sort of emulation service. There does however appear to be a difference in how faculty and members of our node team think about emulation services. Librarians (most of the members of our node are librarians) think of emulation with collections in mind but we’ve started to receive really interesting research requests that don’t follow how we typically do things in our library. These tend to be one off research requests like, "I’m and Phd student researching artists screen savers and there is just one screen saver I need to see running and interact with…not the whole collection. There are many ways of thinking about a service and it is starting to appear as if we have two very different paradymes represented at our institution. The just make it run now and give me access and the more programatic long term preservation activities that we typically engage in as a librarians.

Question 4. There is great interest in our institution for for emulation services. Our challenge has been keeping the stakeholders outside of our node team continually engaged. We’ve been attempting to do this with focused pilots but it is tough to do this for more than one collection at a time. Most of our faculty and librarians need continual technical guidance and we’re already at capacity for being able to provide it which is probably our biggest challenge so far.

I just want to follow on here that Wolbach Library has similarly not performed any kind of “time trials” or measured risk either, and honestly I’m not sure I understand what time trials would be in this context :woman_shrugging:t2:

As for integration with local or cloud-based systems, I think this is an imperative if any service involving EaaSI to scale up.

I think this short “speed blog” I wrote with some colleagues at the Software Sustainability Institute may be relevant in discussions about questions 1 and 2:

I’m not familiar with either of these specifications, but would you consider these in lieu of or in addition to something like CodeMeta? The CodeMeta Project

Hi everyone!

I have to agree with @mgolson with the “emphatic YES” for question two. Creating shared documentation to bounce ideas off one another and develop a “standard” that works for all the different types of institutions represented in SPN and EaaSI nodes would be a good thing I think. I think also creating a hub, or community forum like Michael suggested, would be a good place to plan and start piecing together how we might want to go about it. I also think having it available/update-able after the pilot is over would be beneficial, as it would keep the cohort connected and allow new people to see what has been done, as well see what others have discovered/developed over time or if there’s an easier/better way to do something.

As for question three, some “low level”/informal risk analysis was done when deciding to provide born digital materials via Box to patrons who request it, as well as for instruction purposes. We talked about what file types would work, should we migrate to a more “accessible” version such as PDF for ease of viewing, since everything has to be viewed within Box (we disabled downloads), etc. But all that has been without emulation, so it would be interesting to do time trials or other measurements of emulations both before and after EaaSI being implemented, so we have ways to measure impact. I think with the amount of software used by different writers and artists in the Harry Ransom Center/UT Austin collections, emulation will become a growing focus once we’re more comfortable with it, since certain files are only available to view/interact with using said programs. I would love to see what people do with it and how they interact with it as well.

Interested in reading everyone else’s responses and seeing the discussion!

  1. Documenting and describing resources.

Melanie Swalwell has put out a call to 100s of Australian collecting institutions, asking them to share their records for time-based items (mostly media arts) in their collections. This has netted a number of spreadsheets that the team has normalized. This normalized, searchable database will be available on a page at our new website. This static dataset is a significant addition to researchers in this field, as well as helping us identify large collections that will benefit from preservation through emulation.

The National Library of Australia has provided us a spreadsheet of their software holdings – over 1000 items. They have imaged this software and are using it locally in a VMWare strategy. The National Library has agreed in principle to share these software images with us for EaaSI environment creation.

  1. Melanie Swalwell and Denise de Vries have written a paper on the time taken to image CDROMs and floppy disks. Entitled “Creating Disk Images of Born Digital Content: A Case Study Comparing Success Rates of Institutional Versus Private Collections” https://www.tandfonline.com/doi/abs/10.1080/13614576.2016.1251849?journalCode=rinn20

As for environment creation, I concur with Mifsud’s findings above, that creating an environment that will actually, adequately, and authentically render a complex digital object could take days of effort.

  1. As we continue to educate and build enthusiasm among Australian GLAM institutions and state libraries for EaaSI, issues of training around disk imaging and environment creation and use are becoming more imperative. We imagine we will need some sort of forum like this one, plus some sort of project management to support training and troubleshooting the deployment of EaaSI. The combination of the innovativeness of EaaSI brings with it the strong need to maintain the community around the service.
1 Like

This whole thread is amazing but just jumping in to quickly note: @Cynde_Moya If you can get the vmware images from NLA also they should be directly importable into EaaSI, potentially saving a lot of configuration time.

WOW! @us-hosted members!

THIS IS SUCH. A. RICH. THREAD!!!
Thank you all so much for your thoughtful responses!

As a follow-on to the post about the EaaSI Monthly Node Meeting - the EaaSI team will parse this thread in more detail in order to structure some break out discussions for our April Monthly Node Meeting - and in the meantime, we will follow up on some of the specific resources, suggestions, and reflections that many of you have shared.

For those of you that have not had the chance to respond yet - feel free to reply to your colleagues comments and keep the conversation going, especially if there are efforts being described that intersect with yours or your organization’s thinking/action.

More discussion to come!