Open Access to Research Data

Intro

The research data management cycle

The key takeaways from this article are

1

Research data and their independent publication are becoming increasingly important.

2

The FAIR principles are a de facto standard for research data.

3

Organisational, legal, and infrastructural obstacles before publication can be overcome. 

Research Data

Scientific findings in text form are based, as a rule, on research data. Research data come in a wide variety of forms and types. They comprise all (digital) data generated during the scientific process, for example, through measurements, simulations, interviews, source work, or code and software. The management of these research data has become an essential element of good research practice and is anchored in policies at many institutions. Whereas in the past, research data were often neglected as a mere accessory to publications, made available as a matter of form, or provided only upon request, a strong trend towards the independent and prominent publication of research data in an open format is apparent. In some disciplines, the publication of data is on its way to becoming the main outcome of scientific work.

Reasons for Publishing Research Data

Research data enable the replication and transparency of scientific results. They are thus the basis for transparent science. Reusability promotes the reanalysis of data, the merging of data from different sources, and thus the opportunity to conduct further research with existing data and to generate new knowledge. Ideally, reusability includes the right to download, copy, disseminate, and automatically process the data and to use them without financial, technical, or legal restrictions. The publication of research data enables their citability and therefore contributes to the scientific reputation of the authors. Through research assessment reforms driven by initiatives such as the Coalition for Open Research Assessment (CoARA), research data are increasingly recognised as a scholarly achievement in their own right.

Positions and Drivers

In 2016, the European Union (EU) integrated the Open Research Data (ORD) Pilot into the funding programme Horizon 2020. The ORD Pilot provided for the publication of research data according to the premise

"as open as possible, as closed as necessary".

Participation was voluntary. In Horizon Europe, the successor programme to Horizon 2020, which runs until 2027, open science is designated as the modus operandi and open access to text and data publications and the provision of the data according to the FAIR principles are mandatory.

The German Research Foundation (Deutsche Forschungsgemeinschaft; DFG) refers to the provision of data according to the FAIR principles both in its Guidelines for Safeguarding Good Research Practice and in its separately published Guidelines on the Handling of Research Data. When submitting proposals, applicants are required to provide statements on research data management and data management plans. These statements are taken into account in the evaluation. The German Federal Ministry of Research, Technology and Space (BMFTR) and research funding foundations, for example, also require details of the reuse and exploitation of the data. The German National Research Data Infrastructure (NFDI), which is jointly financed by the Federal Government and the Länder, has proved to be a further driver. According to its self-descriptionthrough the work in 26 consortia, 

"NFDI systematically indexes and networks valuable scientific and research data for the entire German science system and makes it available for sustainable and qualitative use.”

The Austrian Science Fund (FWF) introduced an Open-Access Policy for Research Data in 2019, in which it requires its grant recipients to draw up a data management plan (DMP) and to provide open access to the research data “underlying the project’s academic publications”. The Vienna Science and Technology Fund (WWTF) does not mandate open access to research data in its Open Science Policy. However, it “strongly encourages grantees to provide access to shareable research data”. The Austrian Open Science Policy, which came into force in 2022, aims inter alia to ensure that research data are published according to the FAIR principles. Moreover, several Austrian universities have adopted a research data management policy or guideline.

Further information (in German) can be found on the Austria pages of forschungsdaten.info.

The Swiss National Science Foundation (SNSF) also considers open access to research data to be a significant contribution. Further information (in English) can be found on the Switzerland pages of forschungsdaten.info.

As in the case of research literature, one argument for making research data available in open access is that their production was financed with public funds. At an early stage in the history of open access, the Berlin Declaration on Open Access to Knowledge in the Sciences and Humanities recognised data as objects that should be made openly available. Besides the intrinsic motivation to be able to work more efficiently in increasingly data-driven research with the help of good data management and to benefit from open data oneself, the main drivers of the publication of data are the research funders.

FAIR Principles

In their guidelines and policies, various national and international research funders, for example the EU and the German Research Foundation (DFG), aim to encourage compliance with the FAIR principles. The German National Research Data Infrastructure (NFDI) has set itself the objective of making data “FAIRfügbar”. This is a play on the German word verfügbar, which means “available”. The acronym FAIR stands for Findable, Accessible, Interoperable, and Reusable. The term FAIR was coined by the FORCE11-Community and published in the journal Scientific Data on 15 March 2016 (Wilkinson et al., 2016). Support for the FAIR principles can be found inter alia in the G20 Leaders’ Communiqué issued at the end of the Hangzhou Summit in 2016. The FAIR principles are an internationally recognised standard for the handling of research data. “FAIR data” does not necessarily mean that the data are openly available. 

The four individual elements of FAIR mean:

  • Findable: For the data to be reusable, they must be easily findable. To render them findable, the data are described with rich human- and machine-readable metadata.
  • Accessible: Access to the data found must be possible according to clear rules; authentication and authorisation must be defined.
  • Interoperable: To use data and integrate them with other data, an accessible, shared, and broadly applicable language is needed for knowledge representation. Metadata use standardised vocabularies.
  • Reusable: The description of the data and metadata facilitates their use in different contexts. Suitable data licences are used, and the data meet domain-relevant community standards.

The FAIR principles are complemented by the CARE Principles for Indigenous Data Governance. The acronym “CARE” stands for Collective Benefit, Authority to Control, Responsibility and Ethics. On the basis of the CARE principles, researchers are made aware of the need to uphold the rights and interests of indigenous communities as part of efforts to promote open data and open science – irrespective of whether their research focuses specifically on the communities themselves or otherwise affects them.

For research software, the FAIR principles are complemented by the FAIR Principles for Research Software (FAIR4RS Principles).

In connection with the AI transformation, the acronym “FAIR” may also stand for Fully AI Ready. It means that data processed according to the FAIR principles are mostly machine readable and can therefore be processed by AI.

Publication

When publishing research data, a suitable repository should be chosen – where possible one that provides open access to the data. A disciplinary repository that is well established in the community in question should always be the preferred choice, as one’s own data are thus in a good, specialised context, and findability is easier. The Registry of Research Data Repositories, re3data, can be used to select a suitable data repository. If no suitable disciplinary repository can be found, general or institutional repositories can be used. Well-known examples of general repositories are Zenodo and Dataverse.

To guarantee the long-term provision and findability of the data, a permanent address must be assigned. This persistent identifier also ensures the citability of the datasets. Preferred identifiers are the Digital Object Identifiers (DOIs) provided by the DataCite consortium.

As the reusability of research data is greatly limited if they lack adequate descriptions and metadata, it is imperative that they be curated in accordance with the FAIR principles before publication. Metadata should be assigned at the earliest possible point in time during the research process. They comprise both technical metadata (e.g.: When and by whom was the dataset collected?) and substantive metadata (e.g.: What is the content of the individual variables?). There are specific metadata standards for numerous disciplines. These standards are promoted, maintained, and published in particular by the NFDI consortia (see the RDA Metadata Standards Catalog). During curation, the data are, above all, technically checked. This includes checking the data format, the basic access, and the formal accuracy. Regarding data format, long-term accessible and open data formats should be used, for example. The checking of the data content must be carried out mainly by the researchers themselves. The curation ends with the choice of a suitable licence. The Creative Commons (CC) licences have proved their worth (more information can be found on the English-language pages of the research data portal forschungsdaten.info).  However, CC licences are not suitable for licensing code or software. There are special open source licences for this purpose. 

Practical tip

Tips and tricks on how to make research data openly available to the community are available here in the slides for the Open Access Talk Research Data & Open Access - How to Publish Your Data in German.

Challenges

Three important challenges in research data management (RDM) and in the publication process should be mentioned as surmountable obstacles:

Organisational

Research data management and curation require additional competencies. New professions, such as data curator, data steward, and data scientist, are emerging, and research institutions and infrastructure facilities must make corresponding resources available and provide their employees with the necessary basic and continuing training. Research funders now finance data stewards, for example in collaborative research centres or excellence clusters. Institutions must meet the increased demands by implementing policies and adapted procedures.

Legal

Research data may be personal and very sensitive. These data must be anonymised before publication, or access to them must be restricted in such a way that no data protection rights are violated. Copyright aspects in connection with research data should not be ignored. This applies both to the publication of one’s own data and to the reuse of existing data. It is imperative to obtain legal advice at an early stage in the research process.

Infrastructure

Especially in the natural sciences, the amounts of data generated can quickly become very large. Past experience shows that the volume of data will steadily grow. Handling petabyte-scale data places demands on storage, backup, archiving, and transfer. 

Reservations

One criticism frequently voiced by researchers stems from their concern that others will benefit excessively from their wealth of data, and that they will not achieve the reputation they need for their scientific careers. It should be made clear in this connection that publication should be sought as early and as comprehensively as possible. However, the data may still be published at a later point in time – after the analysis or subject to an embargo period. This reservation is also being addressed through proposed reforms in the assessment of research. Sovereignty over the data can remain with the data producer. The additional costs associated with the required handling of the data are also frequently mentioned. It is already possible to also request funding for these costs when applying for research funding. Data management must be regarded as a key element of scientific research and be adequately staffed and funded.

Outlook

The process observed in recent years will further intensify. The publication, reuse, and linking of research data have become the scientific standard. Many journals now require compliance with data availability statements upon submission. Furthermore, dedicated data and software journals are establishing themselves as additional publication venues for research data alongside repositories.

Overall, it is to be expected that the boundaries between open access, text publications, and research data will become increasingly blurred, that the topics and tasks will overlap, and that a change in mindset will be brought about in the scientific community under the umbrella term “open science”. For acceptance in the scientific community, data and their handling must be recognised as a scholarly achievement.

References

  • Wilkinson, M. D., Dumontier, M., Aalbersberg, Ij. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.-W., Bonino da Silva Santos, L., Bourne, P. E., Bouwman, J., Brookes, A. J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C. T., Finkers, R., … Mons, B. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3. https://doi.org/10.1038/sdata.2016.18

Further Reading

Further Links

Content editor of this page: Matthias Landwehr, University of Konstanz (Last updated: May 2026).
Special thanks are due to Anna-Karina Renziehausen, TIB – Leibniz Information Centre for Science and Technology and University Library.