The repo_check module finds links to research repositories in a paper, retrieves the list of files in each, and reports what is shared. It is the foundation for every other module that reads repository files: code_check and data_check both build on the file list repo_check produces. See Repositories it searches below for the full list of platforms it recognises.
Beyond listing files, repo_check runs a set of checks on what it finds:
whether each repository has a README;
whether files are bundled into archives rather than shared individually;
whether any files are unclassifiable by name or extension;
#>
#> - We found 1 empty repository.
#> - We found 78 files in 4 repositories.
#> - We found 6 README files and 3 repositories without READMEs.
#> - We could not classify 1 file.
#> - We found 15 file naming problems to fix.
The table lists every file found across all linked repositories:
repo_check looks for links to the following repository platforms, each via its own link-finding helper and its own file-listing call. Some are a single hosted service (Zenodo, ResearchBox); others are open-source software that many independent organisations each run on their own address, in which case repo_check recognises any address on a maintained list rather than one fixed host — see each row’s note for how many addresses that currently covers.
Platform
Link finder
How files are listed
Coverage
OSF
osf_links()
recursive osf_info(), including registrations (see below)
the Open Science Framework (single service)
GitHub
github_links()
github_tree_files() — one repo-metadata + one full-tree API call, listing the repository in full
github.com (single service)
GitLab
gitlab_links()
gitlab_tree_files()
gitlab.com (single service)
ResearchBox
rbox_links()
rbox_file_download() (downloads and unzips the box)
researchbox.org (single service)
DSpace (legacy)
dspace_links()
psycharchives_file_download() (DSpace’s older REST API; lists without downloading)
PsychArchives and 5 other installations still running this API version
DSpace 7
dspace7_links()
dspace7_file_download() (DSpace 7’s newer REST API; lists without downloading)
64 installations, including Cornell’s eCommons, Georgia Tech’s Digital Repository, and ETH Zürich’s Research Collection
Zenodo
zenodo_links()
zenodo_info()
zenodo.org (single service)
Dataverse
dataverse_links()
dataverse_info() / dataverse_file_download()
121 independent installations, including Harvard Dataverse, DataverseNL, and the DANS Data Stations
Figshare
figshare_links()
figshare_info() / figshare_file_download()
figshare.com (any branded subdomain) plus 16 institutions with their own address on the same service, e.g. the University of Leicester and DTU Data
data.4tu.nl (single service; runs Djehuty, a documented copy of Figshare’s own API)
Dryad
dryad_links()
dryad_info()
datadryad.org (single service)
ReShare
reshare_links()
reshare_info()
UK Data Service’s self-deposit repository (single service)
DSpace, Dataverse, and Figshare each recognise more than one address because the same software (or, for Figshare, the same underlying service under a branded address) is run by many organisations. That list is maintained, not exhaustive — an installation not yet on it simply goes unrecognised rather than misidentified, so it is always safe to ask for one to be added.
21.4.1 OSF registrations follow their source project
An OSF registration is a locked snapshot. Many “Open-Ended Registration” entries (common in older schemas) only describe what was shared, with the real files living in the project the registration was made from (registered_from). Since the manuscript usually links only the registration, repo_check follows registered_from automatically:
if the source project is public, its files are listed under the source project’s own URL (not the registration’s);
if the source project is closed or inaccessible, the report flags this explicitly, distinguishing two cases: the registration itself still holds a mirrored copy of the files (nothing is lost, but the permanent home is unreachable), or no copy exists anywhere (the files are genuinely unreachable).
21.4.2 GitHub repositories are listed in full
Every GitHub repository is listed in full, whatever its size, exactly as OSF, Zenodo, ResearchBox, DSpace and local directories have always been treated. Listing a repository is cheap — it costs two API calls and transfers no file contents — so there is nothing to configure and no size threshold to set.
There remain two reasons a GitHub repository can fail to be listed, and neither is one a setting can fix. Either the repository is invalid or inaccessible (it has been renamed, deleted, or made private), or it exceeds GitHub’s own limit of 100,000 items per tree request, in which case GitHub returns a listing it marks as truncated and metacheck reports it as too large to list rather than silently using a partial one. Repositories in either state are recorded in $gated_repos (see below) with the reason, so a report can explain why a paper produced no GitHub files instead of simply reporting none found.
Note that this concerns listing only. How much of a repository is ever downloaded is still capped, by the per-file and per-repository size limits in data_check described in the data_check chapter.
NoteA change from earlier versions
Earlier versions of repo_check took three arguments — github_gate, github_max_repo_size_mb and github_max_files — that skipped listing a repository over a size threshold. These have been removed, so code passing them will now fail. Delete them from any script that used them; no replacement is needed, because listing is no longer gated.
21.5 Study grouping
repo_check runs a preliminary classification and grouping pass over every file it finds:
data_classify_files() labels each file’s data_type (data, code, readme, archive, documentation, output, …) from its name/path/extension;
.data_doc_role() further labels documentation-type files (readme, license, codebook, …);
data_group_llm() assigns each file to a study group, using deterministic path/repo-split rules first and falling back to an LLM (llm_use(TRUE)) only when those are ambiguous — see the LLMs chapter. Root-level files that are collection-wide by nature (a top-level README, ro-crate-metadata.json, or LICENSE) are excluded from grouping, since they don’t belong to any one study.
This pass is preliminary: repo_check never downloads file contents (except where a platform’s own listing call requires it, like ResearchBox), so it classifies from names and paths only. data_check reclassifies afterwards using file contents too, and its classification is authoritative for where a file is placed in a paper’s structure. repo_check’s own pass exists to power its own report — the file-naming check, the classification dropdown, and the study/repository mismatch check below.
21.5.1 Study/repository mismatch
If the manuscript explicitly names studies (e.g. “Study 1”, “Study 2”), repo_check compares that roster against the studies the files actually separate into. A mismatch — a named study with no matching files, or a file-derived group not named in the manuscript — is flagged in the report. A manuscript that never names studies explicitly has no roster to compare against, so this check is silently skipped rather than treated as a mismatch.
21.6 File naming conventions
check_file_naming() checks every file name for machine-parseability: no spaces or special characters, and every file classifiable by name or extension (data_type != "unknown"). Issues are split into two severities:
bad — a real naming problem (should be fixed);
suggestion — not required, but a good habit (e.g. a more descriptive name).
could not be classified by name or extension (data_type is ‘unknown’); add a recognisable keyword (data, code, materials, …) or a known extension
to_err_is_human
Study 1.r
spaces
bad
contains a space
to_err_is_human
Study 1 - csv - dataset 1.csv___CODEBOOK.csv
spaces
bad
contains a space
to_err_is_human
Study 1.csv___CODEBOOK.csv
spaces
bad
contains a space
to_err_is_human
Study 1 - csv - dataset 1.csv
spaces
bad
contains a space
to_err_is_human
21.7 Special file types
21.7.1 Proprietary E-Prime files
E-Prime’s .edat/.edat2 (and .emrg/.emrg2 merge files) are proprietary binary formats that only the E-Prime software can open — metacheck neither downloads nor reads them. The analysable data lives in E-Prime’s plain-text export instead. repo_check flags any .edat/.emrg file that has no matching .txt export alongside it (matched by file stem), so authors know to also share the readable export (File → Export in E-Prime).
21.7.2 Restricted-access files
A DSpace (legacy) item flagged restrictedAccess/embargoedAccess — most often seen on PsychArchives — hides its files behind institutional (SSO) login — they cannot be fetched programmatically. repo_check flags any repository with restricted files, since metacheck cannot examine them and other researchers cannot reuse them without requesting access.
21.7.3 Archives, and looking inside a .zip
An archive shared as a single file hides what is in it. By default repo_check therefore reads each .zip’s own listing and reports the files inside it, in place of the archive as one opaque row. This costs one HTTP range request per archive and downloads none of its contents: a .zip stores a directory of what it holds at the end of the file, and a server that honours range requests will return just that stretch of bytes. The helper that does this is zip_peek(), and you can call it yourself on any archive URL.
Turn it off to leave archives as single rows and make no per-archive request at all:
module_run(paper, "repo_check", peek_zips =FALSE)
This works for .zip only. .7z, .rar and .tar.gz keep no equivalent directory that can be read from the end of the file, so they would have to be downloaded whole, and some formats metacheck cannot read at all. repo_check warns when an archive is in one of these formats, because that limits how discoverable and reusable its contents are. An archive whose listing cannot be read for any other reason — a server that ignores range requests, for instance — is simply left as a single row, and data_check opens it after downloading, as before.
21.8 A clean example and one with problems
The demo paper links to three repositories, and running repo_check on it produces a yellow light — there are problems to address. The summary text reports exactly what:
#>
#> - We found 1 empty repository.
#> - We found 78 files in 4 repositories.
#> - We found 6 README files and 3 repositories without READMEs.
#> - We could not classify 1 file.
#> - We found 15 file naming problems to fix.
Here the issues are a missing README and an archive file whose contents could not be examined. The per-paper summary table counts each file type — note files_readme is 0 (no README was shared) and files_zip is 1 (one archive):
A clean result, by contrast, would include a README in every linked repository (files_readme ≥ 1 for each), share files individually rather than as a zip (files_zip of 0), have no unclassifiable or badly-named files, agree with the manuscript’s study roster, and have no closed registrations — producing a green light with no items to address.
21.9 Options
repo_check accepts arguments for working with local files instead of (or in addition to) online repositories:
# also include files from a local foldermodule_run(paper, "repo_check", local_path ="path/to/downloaded/files")# only check a local folder, skip all online lookupsmodule_run(paper, "repo_check", local_path ="path/to/files", local_only =TRUE)
local_only = TRUE is useful when you cannot or do not want to contact external services — for example when checking a reviewer submission you downloaded as a zip. See the Local Files chapter for details.
model and params are passed through to data_group_llm() for study grouping, and only take effect when llm_use(TRUE) and the deterministic grouping rules cannot place a file:
module_run(paper, "repo_check", model ="gpt-4o-mini", params =list(temperature =0))
21.10 The traffic light and what you get back
Light
Meaning
green
repositories were found; every one has a README; no archive files, unclassifiable files, or bad file names; no study/repository mismatch; and no closed OSF registrations
yellow
at least one repository is missing a README, an archive file was found, files could not be classified, real naming problems were found, the manuscript’s study roster disagrees with the repository’s file grouping, or a registration points at a closed source project
na
no repository links were found in the paper at all
Element
What it contains
$table
one row per file across all repositories: repo_url, file_name, file_path, file_url, file_location, file_size, file_type, data_type, doc_role, group, paper_id
full output of check_file_naming() — one row per issue, with severity
$roster_check
the manuscript-vs-file study roster comparison (roster, missing, extra), or NULL if not applicable
$gated_repos
repositories found but not listed (e.g. a GitHub repo over the size gate, or a private OSF component), with the reason
$group_no_evidence
TRUE if study grouping had no evidence to work from
$report
formatted report sections: repositories, files, README, archives, E-Prime, restricted access, closed registrations, unclassified files, roster mismatch, naming, and a full classification breakdown
$summary_text
plain-text bulleted summary of each check
21.11 Related
The companion code_check module analyses the actual code files that repo_check discovers. data_check re-downloads and reclassifies the data files using their contents (not just names and paths), and runs the data-quality checks on the columns it finds.