Probabilistic research software project sustainability
Creators & Contributors
(please cite this post as https://doi.org/10.59350/zme5j-0jp45)
I previously wrote a blog post, “There's no such thing as sustainable research software,” trying to be somewhat provocative. In it I wrote:
while we can say that research software has been sustained, we can't predict with certainty that it will be sustained in the future, or even fully know what factors we should use to make an uncertain prediction.
However, after thinking about this a bit more, I don’t think this means we shouldn’t try. The three things that have influenced my thinking are 1) a set of CZI EOSS Community Calls, the most recent of which was about Measuring and Assessing Open Source Project Impact and Community Health, 2) probabilistic seismic hazard assessment as I learned about quite a while ago from Phil Maechling at SCEC, where one set of work was to model the ground motion at many locations that would be caused by a large set of possible earthquakes so that the range of possible motion at any one location from the set of possible earthquakes could be predicted, and 3) a group in a recent Dagstuhl workshop that I was part of examined lifecycles and categories of Research Software, starting with research done by Yo Yehudi in her EngD work.
Basically, I think we can try to estimate sustainability (a forward-looking property) as a probabilistic forecast.
If we say that research software sustainability is a measure of the software project’s ability to survive, we can then look at the
- Possible events that might cause it not to survive
- The likelihood of each such event
- The likelihood of the consequence of each event causing the project not to survive
This is similar to a risk registry, where both the probability and impact of the risk are estimated, and the product of the two is used to determine the importance of mitigating the risk, though here all of the “risks” are then summed to determine the project’s sustainability.
For the first item, I think the possible events that might cause a project to not survive are the set of challenge events captured in Towards Defining Lifecycles and Categories of Research Software), such as the main developer leaving the project, a funding grant ending, etc. If we’re missing things there, please let us know.
An alternative view can be found in How Nonprofits Close: Using Narratives to Study Organizations Processes, which focuses on non-profits, with many similarities (and some differences) to open-source software projects. This found that causes for closure of these nonprofits included mission completion, program failure, lack of external commitment (grants, loss of clients), financial difficulties, organizational expansion, organizational contraction, or a lack of internal commitment.
For the second item, to assess each event’s likelihood, we might be able to mine existing repositories to understand their history, perhaps also examining issues or other mentions to understand what challenges have occurred and what their consequences have been. Alternatively, this could be done via qualitative studies, such as interviews or surveys.
And for the last item, to look at the likelihood of the consequence of each event causing the project not to survive, we might be able to use the CHAOSS metrics. For example, there are metrics related to bus factor and funding. We might also need to factor in a metric related to impact, with the idea that if a project has a large impact because it is used by a lot of important research projects or has many other software projects that are dependent on it, these other projects are likely to find a way to sustain the software. (Note that while dependencies in open source can be found via repository mining, usage of research software is much harder to determine, though some recent work on estimating usage of open source exists.)
I don’t think being able to determine probabilistic research software project sustainability will be easy, but I do think it can be done, and I also believe that work along the way will have valuable impacts in both research software and the wider open-source community. Thus, I think that this is worth pursuing.
Acknowledgements: Thanks to Sean Goggins, Beth Duckles, and Dan Sholler for helpful comments on a draft of this post.
Additional details
Description
(please cite this post as https://doi.org/10.59350/zme5j-0jp45) I previously wrote a blog post, "There's no such thing as sustainable research software," trying to be somewhat provocative.
Identifiers
- GUID
- https://danielskatzblog.wordpress.com/?p=1748
- URL
- https://danielskatzblog.wordpress.com/2024/10/29/probabilistic-software-sustainability/
Dates
- Issued
-
2024-10-29T13:03:51
- Updated
-
2024-10-30T21:48:39