There's no such thing as sustainable research software
Creators & Contributors
(with contributions from Michelle Barker, Neil Chue Hong, Matthew Turk, Jeffrey Carver, Hannah Cohoon, and James Howison; please cite this post as https://doi.org/10.59350/naakc-7h373)
Ok, this title is a bit of a teaser, but what I really mean is that while we can say that research software has been sustained, we can’t predict with certainty that it will be sustained in the future, or even fully know what factors we should use to make an uncertain prediction.
As discussed quite a few years ago now, I use “software sustainability” to mean that
the software will continue to be available in the future, on new platforms, meeting new needs
As is the case with many financial products in the US that include a risk warning that’s something like “past performance is no guarantee of future results,” the same is true for software sustainability.
In the financial world, some argue that products that have performed exceptionally in the past are unlikely to do so in the immediate future, because of regression to the mean. Others argue that past performance is based on some underlying factors that will have an effect on future performance. Both of these likely are true to some extent.
I think the predictive aspect is the stronger one for sustainable research software, and that software that has been sustained in the past is likely to be sustained in the future, though of course, unforeseen events (e.g., the key person leaves) can cause this to fail.
In part this is because of a critical mass effect, where research software that is sustained over a long period likely continues because there is a user and supporter community, and this community enables it to better survive many events, as long as enough members of the community are willing to take on new roles to do so.
Of course, as Matt Turk has reminded me, there’s also an aspect of evolution vs revolution (and disruption) that can complicate this. For example, a research software package can be sustained for a while, and it may appear that this will continue, but a different package that is substantially better in some way (a new algorithm, improved performance on new hardware, etc.) will then appear to replace it.
A number of groups, including individuals, projects, and organizations, care about doing as good a job as possible at predicting sustainability, such as developers thinking about the software they work on, funders thinking about the research software they support, users thinking about the research software they use either as end users or within a development process, and researchers who study research software.
It would be very useful if there was a clear set of factors that these groups could use, that would correlate research software that has been sustained with characteristics of that software and its community, including developers and maintainers, users, funders, etc. Neil Chue Hong also suggests that it would be useful to be aware of thresholds that might signal a decreasing probability of future sustainability.
However, while I’ve seen some work trying to do things like this, I don’t think I can point to any one document or website that collects it.
Examples of such work include:
- Analysis of software submitted to and recommended for funding by the CZI EOSS program: see slide 15 of this talk by Dario Taraborelli.
- A preprint of a study of 120 NSF funded projects by Cohoon, Du, and Howison. This found that projects that were already functioning open source projects at the time of funding were much more likely to be sustained in the 4-year period after the grant. They also identified two routes that code took: reorganization (such as code being built in a lab changing to an open source project) or handoff (where the code moved to a different organization entirely). Interviews highlighted these challenges in achieving sustained peer production: high effort in mentoring and upskilling potential contributors, the importance of positioning well within nearby ecosystems of packages, the perception of a much smaller potential contributor pool in scientific software than in more general open source tools (only in the low 10s or 100s).
- Work studying a set of open source projects over a year by Yo Yehudi, Carole Goble, and Caroline Jay.
- Work done by CHAOSS in collecting community health metrics, though this is in the context of open-source software broadly, not research software specifically.
- A review paper on software sustainability in software engineering (more broad than scientific software). They found over 100 definitions, but also 12 papers with metrics.
- 10 simple rules for teaching sustainable software engineering focused primarily within a single lab (but also evolving students into contributors)
- A dataset of NSF funded grants that produced software (which could be used to study sustainability of the efforts) by Brown, Schwartz, Huang, Weber.
There are also qualitative studies of the work needed to maintain scientific software:
- Howison's presentation focusing on the sources of work needed to keep software scientifically useful
- The Work of Making a Software Pipeline Repurposable by Neang, Sutherland, Ribes, and Lee
- A study of the "extra work" needed to move from a personal tool to a sustained project by Trainer, Chaihirunkarn, Kalyanasundaram & Herbsleb.
I would welcome more information along these lines, and I’ll update this post with any responses.
Additional details
Description
(with contributions from Michelle Barker, Neil Chue Hong, Matthew Turk, Jeffrey Carver, Hannah Cohoon, and James Howison;
Identifiers
- GUID
- https://danielskatzblog.wordpress.com/?p=1730
- URL
- https://danielskatzblog.wordpress.com/2024/05/13/no-sustainable-research-software/
Dates
- Issued
-
2024-05-13T18:52:14
- Updated
-
2024-08-31T13:22:24