Scientific Data Management Options
Scientific Data Management: Seven Options, and Where Your Data Actually Lives
A key aspect of the modern scientific method is the requirement for scientists to maintain accurate records. The process as it exists today evolved from the “commonplace” notebooks of the Renaissance into a formal technique used to record observations, describe experiments, test hypotheses, document an accumulated body of empirical evidence, and share the observer’s thought processes and conclusions with a larger audience. The goal of this process is to facilitate verification, promote corrections, encourage collaboration between scientific peers, and support the continuous communication, spread, refinement and extension of collective knowledge. Paper lab notebooks served this purpose well, and over time, customs and conventions related to the format and content of these notebooks arose in the scientific community that encouraged scientists to keep accurate, logical and reproducible notes that could then become the basis for formal publications.
For some historical background and examples, see:
A Brief History of Lab Notebooks:
https://press.asimov.com/articles/lab-notebooks
Erasmus Darwin’s Commonplace Book:
https://www.revolutionaryplayers.org.uk/erasmus-darwins-commonplace-book/
Collection of the scientific works of Isaac Newton:
https://cudl.lib.cam.ac.uk/collections/newton/1
With the digital revolution that began in the late 20th century, the sheer quantity and diversity of data types being collected made paper notebooks untenable, and most scientists have gradually shifted to digital collection, storage and distribution of data. But no single agreed-upon strategy has emerged for the best way to do this, and a bewildering array of tools and strategies have been adopted by individuals and organizations. Each approach answers three important questions differently: Where does the data actually live? How do people create and deposit it? And how (or whether) can it be safely edited later?
Let’s consider the advantages and disadvantages of seven of the most common strategies:
1) Your Own Computer
Even in 2026, many scientists still keep records primarily on their own personal computers, and use them as their sole tool for scientific data management. It’s not hard to see why. Your computer offers the psychological benefit of a sense of control, ownership and privacy, combined with the speed and efficiency that comes from navigating through a virtual space whose organizational schema you know extremely well. Critically, your computer also offers the convenience of outstanding integration with the local desktop applications you probably use every day, allowing you to instantly view and edit almost any file in almost any format. This is so natural to us that we often forget how powerful it is to double-click any file and have it open in a specialized tool specifically designed for working with files of that type. Another plus is that on occasions when you can’t quite remember where assets are, your OS has a moderate ability to search your computer (notwithstanding the fact that, for me personally, the search tools in Windows, SharePoint and Outlook are an endless source of frustration).
Modern operating systems are moderately secure, but ultimately vulnerable — not just because of malicious attacks or the installation of malicious software, but also because of the substantial risk of physical loss, damage or malfunction, especially given that most users, although well intentioned, are not disciplined with their backup routines. Operating systems are also not intended to be used as databases, so in a busy scientific laboratory your operating system would probably need to be supplemented with something like a LIMS or another method for dealing with large quantities of structured, repetitive data. If you have installed database applications, use database-friendly file formats like Excel, or have access to an online database tool, you may be able to cobble together a workable solution based entirely or largely on your own computer.
However, consumer operating systems are not designed to hold files in ways that follow GLP, 21 CFR Part 11 or related compliance rules; there is no concept of an immutable version history or audit trail. Digital signatures could possibly be added through a third-party service, but that turns out to be surprisingly expensive for a relatively small amount of additional functionality. Probably the biggest disadvantage of your own local computer is that it cannot be used in a collaborative manner — and it especially cannot be used to allow and control access to huge numbers of files, or provide versioned, sequential file editing for dozens or hundreds of colleagues in an enterprise environment. Even if you could find a way to let others access your personal computer, there would be no concept of tracking or tracing correct attribution or file provenance to help a downstream audience decipher who contributed what. Another shortcoming is that, out of the box, consumer operating systems do not offer stringent lab notebook functionality (i.e., configurable notebook rules and formatting conventions, templated multi-step workflows, staged approval or signature processes, or any of the other features that were developed for paper notebooks and then transferred to ELNs to allow large research organizations to meet the same legal requirements enforced for paper records).
Where does the data live? With this option, data resides primarily on the hard drive of your computer, although backups of various versions may exist elsewhere depending on your backup strategy.

2) Shared Network Drives or Cloud-Based File Storage (e.g., MS SharePoint)
Use of a shared file repository may help resolve some of the issues of collaborative access to files, but these tools introduce a range of new shortcomings when compared to your own personal computer. Controlling access and permissions can be extremely complex. Viewing or editing specialized files requires a clumsy cycle of downloading files to a local machine in order to use specialized software, then uploading the files back to the server when you are finished — a process that can easily create redundant copies and introduce versioning problems, with users overwriting each other’s versions or unable to discern which version is the latest. Anyone who has attempted to work in a highly collaborative way in SharePoint is familiar with these problems.
Just as with consumer operating systems, there are still problems with versioning, access, permissions and compliance, and a lack of digital signatures. Use of a shared file system can also introduce new issues related to data sovereignty. Where is the data physically stored? Who performs backups? Who has access to it? Who controls that access? Who keeps the hierarchy organized? If the system is permanently connected to the internet, who will ensure, manage and monitor security? Given the rapid post-AI escalation of security threats, can you even be sure that it is possible to keep your data secure in any such system? And once again: none of the legally protective lab notebook features are present.
Where does the data live? With this option, the data exists mainly on the shared drive you have selected — either on your local network or in a known or unknown location in the cloud — but copies may also exist on the computers of the various contributors.

3) Web-Based Scientific Document / Content Management Systems (DMS, EDMS, CMS, SCMS, ECMS)
These systems can have a number of advantages over an off-the-shelf or do-it-yourself shared drive or file store. They tend to have better scientific compliance, may manage versions better than home-grown or off-the-shelf file store solutions, and may do a better job of controlling access and recording provenance and audit information. But if the servers are cloud-based, they still present the same challenges for security, availability, data location and sovereignty, and trust in the vendor that provides the service. Commercial on-premises solutions of this type are disappearing, because they offer a lower margin for the vendor. Costs can be considerable, and again — although these systems are a big improvement over off-the-shelf file storage when compliance matters — they usually don’t include legally protective lab notebook features, because they are intended mainly as file stores, not contemporaneous note-taking platforms.
Where does the data live? With this option, your data typically resides on the cloud or local server where the solution is installed.

4) Laboratory Information Management Systems (LIMS)
A LIMS is typically used as a database of structured data. Most LIMS systems allow you to collect streams of structured data, then sort, filter, annotate, search, and export it, create reports, or perform key calculations and analyses. A LIMS may be used for very specific types of data management and analytical tasks in well-defined environments and workflows, but LIMS don’t do well with collaborative file management or the ambiguity of discovery research. They are not good at creating ad-hoc relationships between different resources or editing different types of files on an as-needed basis. Since they mainly hold raw data, they typically need to be supplemented with some other storage system for the day-to-day editing of files and documents. LIMS also don’t do well with any kind of offline work and cannot always accommodate unexpected edge cases or highly heterogeneous data.
They can, however, provide good collaborative access and control for an enterprise organization that only wants certain people to have access to certain types of well-defined structured data. They are also good for recording data that is continuously or sequentially generated over time or in batch processes, so they are often indispensable in production and manufacturing facilities, or in high-throughput labs performing repetitive analysis or tracking large numbers of samples or batches. Due to the complexity of some of these systems, training and usability can certainly be an issue — but because only a portion of the total functionality is typically used in any given setting, and because LIMS systems are often used in a highly repetitive way, technicians can become power users of the specific, focused, well-defined functionality they rely on every day.
Many high-quality LIMS systems do offer good compliance with rule sets like GLP and 21 CFR Part 11 for the data they manage, but they usually don’t include legally protective lab notebook features — or if they do, unstructured note-taking functionality is offered as a bit of an afterthought. Some LIMS systems are designed to manage physical samples; others may focus on data derived from samples and/or instruments, in which case a separate sample tracking or inventory system may be necessary.
Where does the data live? With this option, your data typically resides on the cloud or local server where the solution is installed.

5) Laboratory Information Systems (LIS)
The differences between a LIMS and a LIS are a bit blurry and vary depending on who you ask, but the most common convention is that “LIS” is more likely to be applied to a structured data collection and management system used in hospitals, biomedical or clinical settings — or anywhere that patient data or HIPAA compliance is likely to be required. A LIS shares many of the same advantages and disadvantages described for LIMS, but the price and deployment complexity can be greater, because the LIS may need additional layers of security and validation — for both the software solution and the host environment — before use, to ensure that all legal requirements for patient privacy are met. Validating systems to assure organizations that they are fit for use with private medical data can be a complex and expensive process, and purchasing space on a pre-validated host designed for this purpose can also be expensive.
Where does the data live? As with a LIMS, your data typically resides on the validated cloud or local server where the solution is installed — with the added requirement that the host environment itself meets patient-privacy standards.

6) Electronic Lab Notebooks (ELN)
ELN systems can resolve many of the collaboration and compliance issues associated with diverse file types, and they allow for the addition of rich explanatory metadata, along with the unanticipated provenance and edge-case details that the other techniques described above cannot easily capture. If the system is compliant with the widely used 21 CFR Part 11 rule set, it will also include a coherent versioning system for files and notes — which in turn can produce a usable, unalterable version history, an audit trail, and contemporaneously recorded unstructured notes that follow the established format and compliance guidelines of research organizations.
However, although some ELNs may include rudimentary LIMS or LIS functionality, working with structured data tends to be a bit of an afterthought in ELNs, just as working with unstructured data is a bit of an afterthought in most dedicated LIMS. An ELN may not be adequate by itself for organizations that are collecting large quantities of highly specialized, repetitive structured data. ELNs are also not usually designed to perform complex analysis of data; rather, they can be used to indicate how and where that analysis was performed, and may include the final conclusions from such analyses.
ELNs tend to be generalist tools, often more flexible than a LIMS or LIS, and increasingly they are designed to act as a central hub for your data — with summaries or reports flowing into the ELN from other sources, or with methods for creating links to additional verbose or structured data stored outside the ELN. If the ELN happens to include a good ability to interact directly with desktop applications on your computer — as is the case with CERF ELN — then the ELN can also reproduce some of the advantages of working with your own operating system, and users can choose their preferred or specialized applications for creating and editing data.
Just as the paper lab notebook was intended as a record of, and a guide to, repeating experiments, procedures and the interpretation of results, so electronic lab notebooks are also focused on telling a human-readable story — one that captures details, indicates where core or peripheral data is stored, and preserves the details of the people, places, procedures and materials used. If users are trained in the proper use of the system, then most or all of the data should be stored only on the ELN server and its associated backups, although some sorts of verbose data may also live elsewhere — hopefully in a location indicated within the annotations of the ELN.
Where does the data live? On the ELN server (and its backups), with a complete version history and audit trail. Bulky external data sets, if any, are linked and described from within the notebook record.

7) Long-Term Repositories
In academia, it is increasingly common — and often required by funding bodies — for data to be deposited into specialized, long-term repositories designed for that task. The objective of these repositories is to make data available to others, and these repositories often assign digital identifiers to data sets in order to make them more compatible with the objectives of FAIR data management. The problem with this approach, though, is that the downstream audience only sees what is in the repository. There is no indication of provenance, and no indication of what the difference might be between the data collected at the bench and the data that was actually selected for deposit into the repository. Connected ELN systems like RSpace can help generate a more authentic record of exactly what was done in the lab, and then send this entire data set directly and seamlessly to a selected repository without modification, creating a clearer picture — but ultimately it is the owner of the data who decides what does and does not enter the repository.
Where does the data live? With this option, a final version of the data is located wherever the repository is installed, but earlier versions, various copies, and rejected data that was not deposited could be almost anywhere — including within any of the other systems mentioned above.

So Which Option Is Right for You?
In practice, most organizations end up with a combination of these tools: perhaps a LIMS for high-throughput structured data, a shared drive that nobody fully trusts, personal computers full of “working copies,” and a repository at the end of the pipeline. The real question is which system acts as the authoritative record — the place where the story of your science is told, where provenance is preserved, and where an auditor, a collaborator, or a future version of yourself can find out what actually happened.
That is the role an ELN is designed to fill. If your work needs to stand up to regulatory scrutiny under GLP or 21 CFR Part 11 — or you simply want a single, trustworthy, versioned home for your team’s records — a compliant ELN such as CERF can act as the central hub that ties the other systems together.
To learn more about CERF ELN, visit https://cerf-notebook.com, or contact the Lab-Ally team at https://lab-ally.com.

Comments are closed.