(RESPONSES DUE Wed, March 24) EaaSI Hosted Pilot Forum Discussion #2

Greetings @us-hosted Cohort Participants!

Hope everyone is doing well this week - being kind to yourselves and finding moments (however brief) during the week to look beyond the screen.

From today through the middle of next week, we are going to engage in our second forum discussion - pausing for a little group reflection before we complete the second part of the software description activity. TIMELINE NOTE: After we complete the second part of our software description activity, we’ll be ready to start FORMAL EaaSI TESTING (around the first two weeks in April)!

Here’s some food for thought to contextualize this week’s forum discussion questions:
Synthesis doc for Software Description Part I

Please respond to the discussion questions below by replying to this post with your thoughts and experiences, any follow-up questions that you want to share, and/or responses to your colleagues’ posts.

  1. In our first discussion, there was close-to-consensus on the value of a coordinated collecting strategy for software as well as manuals. The results of the software description exercise (and the varying levels of metadata provided by the software itself, or by different creators/publishers) re-emphasized the criticality of manuals and other contextual documentation that exists “in the wild” for enabling meaningful reuse of software. Put another way - cultural organizations are only one small part of an already large and still growing software/metadata ecosystem. How do you or your organizations think about our increasing reliance on third party systems within your local systems and workflows? What are the advantages of investing in that broader ecosystem - pointing to third party systems from our local systems rather than re-investing that effort locally? What are the disadvantages? What is your formal or informal criteria for incorporating a third party resource as a dependency in your local systems? As a community, how do we support the longevity of the third party resources that provide the greatest enrichment to our systems and services? (Think: Wikidata, Zenodo, etc.)

  2. In responses to this exercise, it was clear that system requirements are both essential for emulation and an extensive amount of information to capture (with no obvious place in our existing descriptive systems). What are some ways we might automate capture of these requirements, and what would be involved in extending your current metadata application profiles to accommodate discrete fields for each of the major system requirements?

  3. As you’ll see from the last bullet in the synthesis document, there is shared interest in seeing the full EaaSI metadata application profile and mapping that (very granular) set of elements to increasingly less granular standards that put EaaSI in conversation with the broader world of existing library and repository metadata standards. Please share which type of metadata standards you are most interested in mapping EaaSI elements to (e.g., disciplinary standards, content standards, descriptive standards? MODS, RDA, PREMIS, CodeMeta, SMRF, FRBR). Put another way, what are the standards that are of greatest importance to your repository and its users - what does the EaaSI team need to be thinking about in terms of crosswalks that will enable broader engagement with and adoption of EaaSI?

  1. I think this depends a lot of the stability and reliability of the third party system. For instance I love pointing out to Zenodo and being able to support researchers who are doing their own software deposits there. It’s just not scalable or necessary for our library to actually be stewards of all the software artifacts (and their many versions) that are worth archiving, and it would be impossible for us to actively curate everything our community creates. I think the primary drawback here comes down to the fact that we’re pretty much putting this responsibility on software authors who typically have little or no training or time to do this work well so they often make mistakes or omit metadata that would be really valuable to to have. At the very least though the software and metadata that authors do record is archived and made available using a DOI, which means the software is meaningfully citable (authors get credit) and findable. If we were talking about the need to point out to third party systems that are meant for code development (e.g., GitHub) I’m much less enthusiastic. Platforms like this are not meant to act as endpoints for perpetual access to stable software and metadata, so they can’t meet that need well. Software authors often don’t understand this though so they have limited ability to make informed decisions about whether or not it’s necessary to do more than leave their code on a personal website of git hosting platform. In the long run I think information maintainers and software curators should be incorporated into research teams to mitigate these drawbacks.

  2. Have authors populate minimally required emulation metadata (if emulation is what the author would like to enable) in machine-actionable formats (i.e., CodeMeta) and include those files with their deposits when they submit copies of their code to archives. Ideally then archives would be able to create and run validators that could flag files as being poorly formatted or incomplete and make depositors aware of areas where changes/updates are recommended. Individual archives could provide templates customized to their specific systems and eventually work toward making these files indexible and incorporating missing metadata fields (e.g., software version) into their systems’ schemas.

  3. CodeMeta and CFF are going to both be really important in this context more broadly, but when it comes to library-specific systems I think mapping to RDA and taking a look at MARC would be good places to start.

  1. (cont.) Harvard’s catalog also supports MARCXML records, which would probably be the more likely choice for this project instead of MARC21. An extensible markup language would be more future-proof.
  1. I think our institution has a model for coordinating preservation with third-party systems through its involvement in HathiTrust. In many ways, we really do rely on that partnership for preservation of our book-like material. In this case, it’s a consortium that consists of the people who are contributing data – I’m not sure of the structure of Wikidata, for instance, but I suspect it’s a bit different! For folks who use Zenodo, I’m not sure of the ins and outs of their policies, but it does seem like it’s secured through the next two decades. In my dream world, perhaps our involvement with EaaSI could be like our involvement in HathiTrust. (But perhaps someone has some good reasons why we might want to diverge from that mode…

2a. We might need a flow chart for what needs automating. Are we looking to figure out requirements for individual document-like files that we find in our repositories? Scripts or sets of scripts? Published software? How much metadata already exists? I think we might need to enumerate all of the cases where we have a thing that we need to know appropriate environments for, and maybe go from there?

2b. Is anyone here in the pilot part of an organization that is involved with FOLIO development? I’d be up for brainstorming with you whether there something in that space we could explore for extending application profiles to accommodate more granular system requirements.

  1. I am most interested in seeing whether the Common Platform Enumeration Dictionary would be helpful to us, as a community in terms of providing canonical names for software environments. I’m also wondering how useful libvirt elements would be in providing structured descriptions of emulated hardware environments.

I guess if we’re part of the consortium, then maybe HathiTrust isn’t “third-party” so please take that part of my comment with a grain of salt. (Or a large rock of it.)

These are pretty “off the top of my head answers,” so hopefully coherent and helpful!

  1. I think we’re open to incorporating more third party systems into our workflows and systems. We’re [the HRC] currently looking into things such as SNAC and WikiData and how we can increase visibility of our collections. Investing into the broader ecosystem would help manage time priorities for staff by not “reinventing the wheel” and re-doing what has already been done. A disadvantage would be placing all the trust in a system that could be wrong, or go down, or stop being supported, or sunset with little to no advance warning and leave you back at square one with everything. I think having a local backup or something similar would be worth having - not duplicating work exactly, but knowing exactly what is being pointed to and what types of information is there and is considered “essential”. I’m still not 100% sure on policies and how things are incorporated, but we are encouraged to test and play with third party things, and to have them officially added to a workflow, I think a formal proposal needs to be created and submitted. The final question is a great one - I think incorporating them into our processes and making sure they know we’re using them and how, can show their value and possibly help them get further support from other powers.

  2. Is there a way to pull from places that have this information, like the WIkidata QID or some other type of database with information for popular programs? (I say this not knowing if there is such a thing, but there probably is knowing people?) I think we would need to create them from scratch - we could probably crosswalk a few from our existing profiles, but most would need to be beefed up to capture all the required information. The tool a cohort member is working on to extract the information from the software itself seems really cool and I would love to know more about it and test it when possible, just to get a general baseline of information, even if it needs edits.

  3. I agree with Daina and Eric that MARC/MARCXML should be looked at and included, since they are fairly standard. I’m interested in learning more about the resources others have mentioned here though!

Also off the top of my head answers - some of the things mentioned above are new to me

  1. In terms of reliance on third party systems, I think this is becoming the norm, at least within the National Archives of Australia (NAA). However this is a choice made through seeking compatibility for current business systems, as well as attempting to avoid having too much that is “bespoke” in our workflows. Provided the NAA has a decent level of control over what is brought in from third-parties for use, (and that it isn’t necessarily housing collection material or Australian government data on a cloud that is based outside of Australia, and that it meets cyber-security requirements), then we are all fine with third-party. I hope I understand correctly, but in a way we already use (but not yet do we rely on) third party platforms like Pronom to help identify file formats in custody. So from that point of view the advantages are great. I have high hopes EaaSI will embody this ethos as well with the aspect of sharing environments, hopefully reducing the amount of time needed to create an environment. What I have so far found on my local instance is it takes me the greatest amount of time to build the environment so it works optimally. I think I edited my Autoexec.bat and config.sys file over 50 times last week! I am currently rebuilding all my environments for a third time, because it wasn’t clear to me what KVM did, until I went and researched it, by which time I had built and destroyed all my environments twice over. SO I think a criteria for incorporation will definitely need to be a better help file, with greater step-by-step explanations and instructions for pillocks like me who only have a passing understanding of computing.
    I also agree that unless the third party system is stable, reliable and backed by some sort of guarantee it won’t break and lose all of our valuable heritage, the NAA just isn’t going to use it. If it is hard to support with the IT staff we have on hand, then we won’t use it. Frequently, we also like to know we aren’t locked in and as a government agency, we are supposed to go to tender and consider more than one provider. It is a risk if there is only one provider in the market for a particular function.

  2. Automation will be crucial. The ways to do it unfortunately I cannot say. The hurdle we face at NAA would be convincing an agency to let an automation script run on their systems. Likely they just won’t do it due to concerns it will mess with their business systems. It would need to be approved by the Commonwealth security agencies, and likely would have to be built in a way that it could be trusted. For us, it could be built into a transfer portal hosted by the NAA, and be part of the check-sum pre-ingest process when the SIP (or maybe if we are lucky OPEX) is created but that is something that is still a pipe-dream for us.

  3. The main metadata standards we will have to work with are the Australian Government Record Keeping Metadata Standard and the Commonwealth Records Series System, and then maybe the Archival Control Model (which has some foundations in PREMIS and has been amended to better manage digital representations).

The two main systems we use internally for managing digital collection items are Mediaflex, a commercial AV digital and physical asset management system, and Preservica, recently acquired to replace our in-house developed digital preservation system, and which manages born-digital (non-AV) and digital surrogates (digitised records). Both incorporate extensive software libraries for file characterisation, conversion/transcoding etc, but getting additional libraries incorporated takes a lot of work within the respective community/user groups, and can take years to make happen. We do use 3rd party applications outside those systems, for example for some pre-ingest forensics and characterisation tools, but I think when you’re talking third party systems it sounds like you’re talking about research data repositories for access. In Australia there are quite a few examples such as data.gov.au for open Australian government data, the Australian Data Archive (dataverse.ada.edu.au) hosted by the Australian National University and based on the Dataverse Project and its software, and Research Data Australia (researchdata.edu.au) amongst others. From time to time we’ve contributed to these, especially data.gov.au, but our primary online access system is our own RecordSearch database. Unfortunately RecordSearch is very limited in its ability to provide useful access to research data - in fact it’s pretty much non existent ; ) But this is something we’re working on in a discovery layer project, but at the moment it’s not really looking at meaningful access to datasets. That said, independent of, but happening alongside, the EaaSI pilot, the area of NAA Tim and I work in is leading a project on selection, transfer, and preservation of databases/datasets, which is looking at what components of government agency business systems we need to transfer into our custody for to provide the best access to data down the track, eg some combination of SQL dumps, disk images (for future emulation capability), raw data in CSV, SIARD files created using the Database Preservation Toolkit (DBPTK), as well as all necessary technical documentation to make sense of the data. Up to this point we haven’t had a consistent approach, mostly we’ve received CSV exports, plain txt files, and even SQL dumps (that is, the native SQL backup format for the database management system). Providing public access in the future may well involve some third party system or combination of systems, but this is yet to be worked out, and may take a while! EaaSI may well be an option, but we could also make data available through the Database Preservation Toolkit if that proves viable, or one or more of open data repositories mentioned. I can see us trying to figure out best access options at the point when someone requests a dataset (“now how are we going to do this?”)

  1. I’m new to Harvard so I’d be interested to hear what other members in the Advisory group thing regarding some subquestions like reliance on third party systems, in particular. Investing in the broader ecosystem of third party systems can have some of the following advantages: more functionality than local system would allow, faster development/customizations, less work in a specific area because it is outsourced. Disadvantages could include: cost/licensing (if any), learning curve/increased training, reliance on third-party or community support/sustenance of system/resource. I’m not sure if we have formal criteria for incorporating third-party resources like WikiData, but for something like SNAC, the criteria revolve around: highlight relationships between entities and archival collections; reveal entities’ role in the process of archival collecting; focus on underrepresented/marginalized entities. Similar criteria might follow for WikiData. For example, there’s a project going on at Harvard to enhance WikiData entries corresponding to the Arthur Freedman collection. Similar criteria would be extended to collections including software or assets dependent on “old” or “obsolete” software.

  2. I don’t have a complete answer for this question. My nascent thinking on this subject is that it would be great if we could pull data from WikiData or WikiBase to automate capture of these requirements. As for extending our current metadata application profiles to accommodate these discrete fields, if we were to extend MARC (like the other Harvard folks are talking about) we could maybe look at using METS in the MARC Description field. Not super elegant. PREMIS could in theory be extended to include QID’s for example (I’m a little rusty but I think this is accurate) but doesn’t need much extension imo to accommodate system requirements because it’s suited to capture this sort of metadata.

  3. I’m very interested in EaaSI being mapped to PREMIS. PREMIS has many semantic units which aim to describe the original rendering environment and relationships between software stacks that would map well to EaaSI elements. Others mentioned MARC, which I’m curious about, but it makes me wonder why MODS wouldn’t be suitable then? MODS has an extension element that could be used.

How do you or your organizations think about our increasing reliance on third party systems within your local systems and workflows?

For me, these questions arise in the context of supporting software and data-intensive research at Penn State. I help manage Penn State’s open access research repository, ScholarSphere, which is developed in-house. Since 2012, when ScholarSphere was launched, the landscape of repository (and repository-adjacent) technologies has grown tremendously, and the necessity of developing such a platform within the library is less and less clear-cut. Continuing this kind of project means being strategic about its role and complimentarity to third party systems. I actually think projects like Zenodo can contribute to the longevity of in-house systems like ours by raising awareness around common goals – a “rising tide lifts all boats” sort of argument. They also provide a useful stock of user interface patterns and functionalities to build-on and refine, which eases development. I guess the point I’m coming to in writing this is that even in the context of “in-house” systems, we rely on third party systems to help entrench concepts, user patterns, practices, and conventions. (I’ll admit, though, that researchers’ preconceptions of what a “repository” is can be an obstacle sometimes.)

Please share which type of metadata standards you are most interested in mapping EaaSI elements to (e.g., disciplinary standards, content standards, descriptive standards? MODS, RDA, PREMIS, CodeMeta, SMRF, FRBR).

Thinking super practically, we’d be interested in figuring out how EaaSI metadata elements fit into or complement the metadata landscapes of ArchivesSpace, DSpace, and Archivematica.

This is kind of out of left field, but it’s on my mind from other work we’re doing, so I’ll throw it out there – in the context of software created by academics for research projects, being able to map to the UN Sustainable Development Goals would be interesting.

Greeting @us-hosted cohort members!

Thank you all so much for your thoughtful and instructive replies to our second set of forum discussion questions! The team will continue reviewing responses and we’ll be sure to incorporate aspects of the discussion into our April EaaSI Monthly Node call (reminder with date, time, and call information is forthcoming).

Tomorrow, I’ll post the instructions for Part II of our Software Description activity. After Software Description wraps up, we’ll have one more forum discussion before launching into Formal Testing. For Formal Testing, each of your node teams will receive a log-in to their own cloud-hosted instance of EaaSI. The node lead designated for each team will be able to create additional accounts and logins to each instance. Ethan will publish instructions associated with a Formal Testing protocol comprised of a set of specific tasks/actions that we ask all node team testers to complete, with ample room for notes and feedback on challenges each user encounters along the way.

We know it’s been a busy first quarter of 2021 - and we want to express our gratitude and appreciation for each of you. Stay tuned for additional forum posts today and tomorrow!