Published August 8, 2025 | https://doi.org/10.59350/gigablog.6234

Help! I need curation. ISMB meets the Beatles.

Creators & Contributors

Birthdays, BOSC, Beatles and Bioinformatics with a Merseybeat

Conference season is upon us, and the GigaScience team have just returned from a magical mystery tour to Liverpool. Regular readers will know GigaScience launched at the ISMB (International Conference on Intelligent Systems for Molecular Biology) in 2012, and every year we attend and celebrate our birthday at the meeting. The beating heart of the scientific program is the many parallel Communities of Special Interest (COSI) meetings, and we have a particular soft spot for BOSC- The Bioinformatics Open Source Conference. This year we were again silver sponsors and enjoyed meeting with old friends and participating in this track.

ISMB2025 in Liverpool
The fab Liverpool dockside close to the conference venue.


They say it's your birthday

This year ISMB2025 was in the Kings Dock on the iconic Liverpool waterfront, and the organisers (including our Ed Board Member Carole Goble) did a fantastic job organising the biggest ISMB ever, with 2225 in person and 300 online attendees. They (and we) embraced the fab Liverpool location, with conference events in the Cavern Club and plates of "scouse" (a hearty local stew that is the derivation for calling Liverpudlians "scousers") handed out at the reception. We celebrated entering our rowdy teenage years with our 13th birthday party at the One O'Clock Gun pub in the historic Albert Dock. This year having a fantastic turnout, and it was great to see many old (and new) friends, authors, reviewers, Board Members, and a number of the conference and COSI chairs. On top of the now traditional GigaPanda cake, our GigaPanda Beatles themed stickers also going down particularly well (see our pictures of the meeting).

Across the Universe (of biomolecular interactions)

As with the field of bioinformatics in general, one trend over the last few years has been increasing focus on AI and Machine Learning methods, and this year it really jumped out of sessions in the COSI tracks to dominate the meeting in general. The opening keynote setting the scene, with 2024 Nobel Prize in Chemistry winner (and GigaScience author) John Jumper presenting his groundbreaking work with Google Deepmind that has revolutionized protein 3D structure prediction with Alphafold. Giving some first hand insight on how they created and built Alphafold leveraging the CASP Critical Assessment challenges. These Critical Assessment competitions coincidentally linking with many of the ISMB COSI communities such CAMDA and Function (that in it's previous AFP guise we previously published a series). The other keynote talks were also very heavy on AI, with James Zhou presenting on computational biology in the age of AI (agentic) agents, and Charlotte Deane talking about AI-driven structure-based drug discovery. Amos Bairoch's keynote on the past, present and future of biocuration could not avoid the topic, but he provided many caveats and warnings that AI won't work properly without well curated training datasets. And how key human biocurators and biocuration are to this.

Everybody's Got Something to Hide (Except Me and My Open Science)

BOSC covered many of these same themes and warnings, the keynote from Chris Mungall on Open Knowledge Bases in the age of generative AI providing his first hand experience on the advantages that Agentic AI applications and AI coding agents can provide to biocuration. In the AI/ML track Gavin Farrell  presented the DOME Registry we recently published, and showcased GigaScience's integration of the DOME DSW wizard and registry into our peer review processes. Last year Chris Armit from the team presented this at BOSC, and we recently posted the video of a follow-up talk a few months later at the ICG meeting. Alongside the risks and challenges from trying to scale things with these approaches. On top of AI, BOSC continue to cover our old favourite topics like computational workflows, and reproducibility platforms and projects. Jose Espinosa-Carrasco presented on nextflow and the associated nf-core community, showing data that nextflow has this year become the most widely used workflow management system in bioinformatics. Phil Reed presented on the RO-Crate packaging system for capturing FAIR research outputs (showcased in our FAIR Digital Object paper), which was timely as we have a summer intern working on integrating this into GigaDB. BOSC ended with the traditional panel, this year on the topic of Data Sustainability and featuring our Editor in Chief Scott Edmunds alongside Tony Burdett, Nicky Mulder, Varsha Khodiyar, Chris Mungall, and expertly moderated by our Ed Board Member Monica Munoz-Torres. This brought together a variety of perspectives to explore the challenges and possible solutions for achieving data sustainability, but it was hard to find concrete solutions in this time of funding cuts and long established data repositories closing down.

The BOSC Data Sustainability Panel featuring our EiC Scott (picture from BOSC)

Other tracks and COSI's followed many similar themes, the NIH having a specific track on GenAI, Cyberinfrastructure, Digital Twins, and Quantum Computing. And the NIH Office of Data Science Strategy having a joint session with ELIXIR also covering data sustainability and the challenges of running research data infrastructure. CAMDA amongst its many challenges this year tackled Electronic Health Records, with a Health Privacy Challenge and panel trying to address health data security issues through privacy preservation and synthetic data approaches. With the unwieldy size of this years meeting (and the poster sessions being intimidatingly big in the giant Arena section of the ACC) the ISCB has been doing a big push on internationalizing the society and meetings, and this year a new ISCB-China grouping was welcomed to the fold and also had a workshop added to the scientific program. Watch this space for future ISCB-China meetings, alongside the upcoming ISCB-Asia/GIW meeting hosted in our Hong Kong home this December.

Paperback Writer

Since last year a new feature of ISMB is the Publications track, and while last year we participated in the program (see the video of Scott's John Candy themed talk on the challenges with reviewers) this year a different bunch of editors and journals covered the topic of "Navigating Journal Submissions". While it was a new panel and array of topics, many of the same issues as last year kept being brought up in the author questions showing that publishing is a topic that has required this type of regular forum. AI and Machine Learning was again discussed in the session. And like us, Michael Sternberg discussed seeing a huge rise in Machine Learning papers in the journals he edits (JMB and NAR). One concern from these is the issue of data leakage, as there mustn't be the same protein in the training and testing (and some journals require holdback sets to address this). Shaking things up a bit, Michael Markie from eLife ended the session presenting on the novel eLife "Publish, Review, Curate" model that our GigaByte is also touching into through Sciety integration.

Hello, Goodbye

While the topic of AI dominated ISMB2025, many of the speakers kept coming back to the need for the human-in-the-loop. The essential part that human curation plays in training and testing of AIs demonstrating this. The meeting was book-ended on the last day with John Jumper's co-Nobel awardee David Baker doing a virtual talk to also endorse the part CASP played in his and John's success, which has been particularly topical with the NIH recently pulling support for the challenge and Google having to offer bridging funding. David saying despite the incredible rise of AI, there will be a role for human scientists by 2045. Leaving with this reassuring note we human members of the GigaScience team look forward to many more in-person ISMB meetings and birthdays to come.

References

Varadi M (and Jumper J), et al., 3D-Beacons: decreasing the gap between protein sequences and structures through a federated network of protein structure data resources. GigaScience. 2022 Nov 30;11:giac118. https://doi.org/10.1093/gigascience/giac118

Attafi OA et al., DOME Registry: implementing community-wide recommendations for reporting supervised machine learning in biology. Gigascience. 2024 Jan 2;13:giae094. https://doi.org/10.1093/gigascience/giae094

Niehues A et al. A multi-omics data analysis workflow packaged as a FAIR Digital Object. GigaScience. 2024 Jan 2;13:giad115. https://doi.org/10.1093/gigascience/giad115

Additional details

Description

Birthdays, BOSC, Beatles and Bioinformatics with a Merseybeat Conference season is upon us, and the GigaScience team have just returned from a magical mystery tour to Liverpool. Regular readers will know GigaScience launched at the ISMB (International Conference on Intelligent Systems for Molecular Biology) in 2012, and every year we attend and celebrate our birthday at the meeting.

Identifiers

UUID
7d573fa6-35ed-490e-8ff9-c163a18c3853
GUID
https://gigasciencejournal.com/blog/?p=6234
URL
https://rogue-scholar.org/records/f8wvg-8m429

Dates

Issued
2025-08-08T05:49:20
Updated
2025-08-08T05:49:20