Static PDFs Hinder Scientific Progress -- Composable Research Offers Solution
The PDF's Shadow: Unpacking the Hidden Costs of Static Science and the Dawn of Composable Research
This conversation with Rowan Cockett, CEO of CurveNote and co-founder of the Continuous Science Foundation, reveals a critical, often overlooked crisis in scientific research: the inherent limitations of static, PDF-based communication. Beyond the obvious challenges of data access and code replication, the core thesis here is that our current publishing paradigms actively hinder scientific progress by creating friction, obscuring value, and failing to incentivize the very composability that drives innovation. The hidden consequence? A slow, inefficient scientific engine where brilliant discoveries remain locked in inaccessible formats, delaying future breakthroughs. Anyone involved in research -- from students and academics to funders and publishers -- who seeks to accelerate discovery, improve collaboration, and gain a competitive edge in scientific advancement will find immense value in understanding these systemic flaws and the emerging solutions.
The Obvious Problem, The Deeper Flaw: Why PDFs Are a Drag on Discovery
The scientific community grapples with a reproducibility crisis, a problem often framed around missing data or incompatible software environments. While these are significant hurdles, Rowan Cockett's insights point to a deeper, more insidious issue: the pervasive, static nature of scientific communication, epitomized by the PDF. This isn't just about file formats; it's about a fundamental disconnect between how science is done -- computationally, iteratively, and collaboratively -- and how it's communicated. The immediate payoff of a static PDF is its familiarity and ease of printing, but this convenience masks a cascade of downstream effects that stifle progress and create hidden costs.
The sheer volume and complexity of modern scientific data have outpaced the capabilities of traditional communication methods. What once fit on a page now spans terabytes, processed through intricate pipelines. Yet, we often attempt to convey this complexity through static images and links to disconnected datasets. This disconnect isn't merely an inconvenience; it's a barrier to entry. Researchers attempting to build upon existing work face an arduous month-long process of piecing together code, data, and narrative, a task that could be drastically simplified.
"The scientific systems for communication just have not kept up to that at all. So we need better ways to share research, better ways to incentivize the more modular sharing and continuous sharing of research as well."
This friction is exacerbated by the "publish or perish" culture and the project-based nature of grant funding. These social and economic structures often disincentivize the long-term maintenance and accessibility of research outputs. The immediate need to demonstrate progress for the next grant or promotion overshadows the slower, more arduous work of creating truly reusable scientific assets. The result is a system where valuable computational work, the backbone of many scientific endeavors, goes uncredited and unrewarded.
The Hidden Cost of "Good Enough": When Familiarity Breeds Stagnation
The scientific publishing ecosystem, dominated by established journals and societies, often exhibits a form of "complacency." Their current business models, built around static PDFs and traditional review processes, are profitable enough that the incentive to invest in radical change is weak. This inertia creates a significant bottleneck. While preprint servers like arXiv and BioRxiv offer a faster route to dissemination, they often inherit the same static communication paradigms. The inherent limitations of these formats mean that even when data and code are shared, they are often uncurated, lacking the necessary context for true reuse.
Consider the example of microscopy data: a terabyte-and-a-half dataset that, in a modern context, could be explored interactively, much like Google Maps. Today, however, scientists often resort to screenshots within static papers, providing only a limited advertisement of the underlying data. This is a prime example of how an immediate, familiar solution (the PDF) creates a long-term disadvantage by disconnecting the narrative from the executable science.
"There are excellent tools out there like Jupyter notebooks and other tools that have that sort of more literate programming style to them, and those are the types of ideas that we want to bring into the sharing of scientific research, as well as promoting better ways to put your data into sort of more accessible spaces and formats."
The disparity between commercial data management tools and the bespoke formats common in research (like HDF5 or GeoJSON) further complicates matters. While open-source solutions and open data standards are emerging as crucial bridges, the integration layer -- connecting data, code, and narrative into a cohesive, explorable whole -- remains a significant challenge. This is where platforms like CurveNote aim to intervene, by providing a scientific content management system designed for computational research, bridging the gap between how science is conducted and how it's published.
The 18-Month Payoff: Building Moats with Computational Narratives
The true competitive advantage in science, as highlighted by Cockett's work, lies in embracing the computational and iterative nature of modern research and translating that into accessible, reusable formats. This isn't about simply sharing code or data; it's about creating "computational narratives" -- interactive articles where results can be reproduced on demand. Imagine clicking a play button on an image in a scientific paper and having a cloud environment spin up, execute the code, and reproduce the finding. This capability drastically lowers the barrier to entry for others to build upon that research, moving from a month of painstaking reconstruction to a simple click.
This shift from static to dynamic communication offers a delayed but significant payoff. While the immediate effort to create these interactive narratives might seem daunting, especially when journals are still accustomed to PDFs, the long-term benefits are substantial. It fosters a true "standing on the shoulders of giants" ethos, accelerating the pace of discovery. Furthermore, this approach inherently promotes better data management practices. By necessity, building these interactive narratives requires curated, accessible data and well-maintained code, addressing the very issues that plague reproducibility.
"If you can jump into somebody's research stack with a click of a button, that suddenly you're going from a month of work of sort of reading into their papers, finding their GitHub repository, digging out their data set on Zenodo or some other service, and matching it all together in your space, as well as installing their environments and libraries. If that bar comes to zero, then just the possibility to go in, tweak some parameters, have a look at the code, I think the ability to build on somebody else's work, reuse it, that is the space that is like standing on the shoulders of giants, and that is sort of the ethos of science that we're leaning into hard."
This focus on composability -- the ability to integrate modular components of research -- is the next frontier. Unlike software, where packages can be readily combined, scientific research lacks this inherent composability. Creating standards and tools that enable this will unlock multiplicative progress, allowing scientists to build upon each other's work in ways currently unimaginable. The challenge, then, is not just technical, but socio-technical, requiring community-wide adoption and new standards, like the proposed Open Exchange Architecture (OXA), to move beyond the ingrained paper-based mentality.
Key Action Items
-
Immediate Action (Next Quarter):
- Educate Yourself on Literate Programming: Explore tools like Jupyter Notebooks and MyST Markdown to understand how code, data, and narrative can be integrated.
- Evaluate Current Data Sharing Practices: For any research or data analysis, document the steps required for someone else to reproduce your work from scratch. Identify friction points.
- Advocate for Contextual Data: When sharing data, prioritize providing clear README files, code for processing, and environment specifications alongside the dataset itself.
-
Short-Term Investment (Next 6-12 Months):
- Experiment with Computational Notebooks for Reports: For internal reports or smaller analyses, try generating them using Jupyter Notebooks or similar tools to embed executable code and results.
- Explore Cloud-Optimized Data Formats: Investigate formats like Zarr for storing and accessing large datasets, especially if cloud-based storage is utilized.
- Engage with Open Science Communities: Participate in discussions or working groups related to reproducible research, open standards, and scientific communication.
-
Long-Term Investment (12-18 Months+):
- Champion Interactive Publications: For published work, explore platforms like CurveNote or Jupyter Book that support interactive and computationally reproducible articles. This requires patience as the ecosystem evolves.
- Contribute to Open Standards: Support initiatives like the Open Exchange Architecture (OXA) that aim to create new, modular standards for scientific communication.
- Re-evaluate Credit and Incentive Structures: Advocate within your institution or field for recognizing and rewarding the creation of reusable research assets, not just traditional publications. This requires sustained effort to shift cultural norms.