metacheck_cache_info()6 Caching and Reuse
The two slowest and most expensive parts of a Metacheck run are querying a language model and downloading files from an online repository. Both are re-encountered constantly: you re-run a report after tweaking a module, you process the same paper in several modules, or you build an archive from hundreds of repositories overnight and it stops halfway. To avoid paying for the same work twice, Metacheck keeps two on-disk caches — one for LLM responses, one for downloaded repository files — plus a small in-memory cache used by a single module.
This chapter explains what each cache stores, where it lives, how entries are keyed (i.e. what counts as “the same” request), when it is reused, and how to inspect or clear it.
-
LLM responses are cached automatically (on by default). Re-running a report on the same paper replays the stored answers instead of re-billing the provider. Toggle with
llm_cache(); wipe withllm_cache_clear(). -
Downloaded repository files are not cached by default.
data_checkandcode_checkdownload into a temporary folder that is discarded when the R session ends. Passcache = TRUEto keep the files instead, and clear them withrepo_cache_clear(). - Both caches live in the current working directory, so a project’s cache sits with the project. See where they are with
metacheck_cache_info().
6.1 Where the caches live
Both on-disk caches default to a folder in your current working directory, rather than a hidden per-user location elsewhere on the machine. This is deliberate: these caches are per-analysis rather than shared between projects, so keeping them beside the project makes them easy to find, easy to inspect, and easy to delete, and they disappear when you delete the project folder. A hidden cache is exactly the kind that gets silently lost.
| Cache | Folder (in the working directory) | Contents |
|---|---|---|
| LLM responses | .metacheck_llm_cache/ |
one .rds per distinct LLM call |
| Repository files | .metacheck_repo_cache/ |
one sub-folder per repository, mirroring its files |
You never need to create these directories by hand; Metacheck creates them on first use. To see where they are and how much space they take, use metacheck_cache_info(), which returns the location and size of both:
Because the folder names begin with a dot, they are hidden in most file browsers, and because they sit inside your project you may want to keep them out of version control. Metacheck does not write to your .gitignore for you, so add a line yourself if you use git:
.metacheck_*
You can move both caches somewhere else by setting one option, which is useful if you want a single shared cache across several projects:
options(metacheck.cache.dir = "~/metacheck_caches")Each cache also has its own override for finer control: the metacheck.repo_cache.dir option for the repository files, and the METACHECK_LLM_CACHE_DIR environment variable for the LLM responses.
6.2 The LLM response cache
6.2.1 Why it is safe to cache an LLM answer
Normally you would not cache a language-model call: ask the same question twice and you can get two different answers. Metacheck sidesteps this by forcing temperature 0 on every llm() call. At temperature 0 the model is (as close as providers allow) deterministic: the same inputs produce the same output. That makes the answer reproducible, and a reproducible answer is one you can safely store and replay. The cache is therefore not a shortcut that changes results — replaying a cached answer gives you the same thing a fresh call would have.
This is the same property that makes Metacheck’s LLM use reproducible in the first place: it is documented alongside the LLM workflow in the Using Large Language Models chapter.
6.2.2 What is stored, and how entries are keyed
The cache is keyed by everything that determines the answer. When llm() is about to make a call, it builds a key from five inputs:
- the model name,
- the system prompt,
- the input text (the row of text being classified),
- the type specification (the structure the answer must take), and
- any params (e.g. a seed).
These are serialised to a canonical byte stream and reduced to an MD5 hash, which becomes the filename (<hash>.rds). Two details make the key robust:
-
paramsare sorted by name before hashing, solist(seed = 1, top_p = 1)andlist(top_p = 1, seed = 1)hit the same entry — list order does not matter. - The type spec is reduced to its printed form, so if you change the requested output structure (add a field, change a description), the key changes and you get a fresh call rather than a stale answer for the old structure.
Change any of these five inputs and the key changes, so the cache misses and a real call is made. This is exactly what you want: a different question is a different question.
Each entry stores not just the parsed answer (the data frame llm() returns) but also the raw provider result and, where the provider exposes it, the model’s reasoning/thinking trace — so nothing is lost by going through the cache.
6.2.3 Turning it on and off
The cache is on by default. Query or set it with llm_cache():
llm_cache() # is the cache on? (TRUE by default)
llm_cache(FALSE) # force every call to hit the provider afresh
llm_cache(TRUE) # re-enableDisabling it is mostly useful when you are deliberately testing provider behaviour or measuring latency, and want to be sure no answer is being replayed.
6.2.4 Clearing it
To wipe every stored answer — for example after upgrading a model, or to reclaim disk space — use llm_cache_clear(), which deletes all cache entries and returns how many it removed:
llm_cache_clear() # returns the number of entries deletedYou do not normally need to clear the cache to pick up a changed prompt or model: because those are part of the key, a change produces a fresh call automatically. Clearing is for reclaiming space or forcing a clean slate.
6.2.5 What this means in practice
Re-running a report on the same paper is close to free on the LLM side: the first run populates the cache, and every subsequent run replays it. This is what makes it comfortable to iterate on a module’s non-LLM code, or to regenerate a report, without watching your API bill climb. It also means the llm_max_calls() cap counts actual provider calls — cached replays do not consume it.
6.3 The repository-file download cache
6.3.1 What it does
Modules that read file contents — data_check and code_check — need local copies of the files a repository shares. repo_check only lists the files on OSF, GitHub, Zenodo and PsychArchives (it does not fetch them). When a content-reading module needs a file, it is downloaded once and read from there, and every later read within the same run reuses that copy instead of downloading again.
Whether those copies survive the session is up to you, and by default they do not. With cache = TRUE, the files are kept in the cache folder and reused by later sessions as well, which is what makes a long archive-building run cheap to recover from after an interruption: the files already fetched are still on disk, and resuming skips straight past them.
6.3.2 How entries are keyed
Each repository gets its own sub-folder, named by a filesystem-safe encoding of its URL (the scheme is stripped and any awkward characters are replaced), so two different repositories never collide and the same repository always resolves to the same folder across sessions. Inside that folder, each file is stored at its repository-relative path, not just its basename — so two files both called data.csv in different sub-folders are kept apart rather than overwriting each other.
Reuse is decided by a simple, robust test: if the target path already exists and is non-empty, it is treated as already downloaded and reused. A zero-byte file (the fingerprint of a download that failed partway) is not trusted — it is re-fetched. This means a broken partial download never silently poisons later runs.
6.3.3 Turning it on
Unlike the LLM cache, which is on by default, the download cache is off by default. Both data_check and code_check take a cache argument that is FALSE unless you say otherwise:
# the default: files go to a temporary folder, discarded when R exits
module_run(paper, "data_check")
# keep the files, and reuse them in later sessions
module_run(paper, "data_check", cache = TRUE)With cache = FALSE the files are still downloaded and still read; they are simply written to a session folder that R removes on exit, so nothing accumulates on your disk. Files fetched earlier in the same session are reused either way.
Turn it on when you are checking the same repositories repeatedly, or building an archive over several runs, and the download is worth keeping. Leave it off for a one-off check of somebody else’s repository, where you would otherwise be left with copies of their files you never asked to keep.
Reuse of a cached file is safe because a downloaded file is a byte-for-byte copy of the remote one. What you can control separately is how much gets downloaded in the first place, through the arguments described below.
6.3.4 Controlling what gets downloaded
The download cache stores whatever the modules fetch, and data_check gives you fine control over that:
-
download—"data"(the default) fetches only the machine-readable files the checks analyse (tabular data plus codebook/README files);"all"fetches every file in the repository, which is the right choice when building a complete archive;FALSE/"none"downloads nothing and classifies files by name only. -
skip_types— a vector of file types never to download even underdownload = "all". When left asNULL, the default,"materials"and"unknown"are skipped: stimuli and software, and the classifier’s collection of manuscripts, logs and unrecognised content, which would only inflate an archive. Pass an explicit vector to override this, orcharacter(0)to download every type. The six types are"data","code","documentation","materials","output"and"unknown". Skipped files still appear in the manifest, with the reason. -
peek_zips— look inside each.zipusing an HTTP range request, without downloading the whole archive, and only fetch zips that actually contain data or a codebook, leaving a zip of pure stimuli to be linked instead. It isFALSEby default indata_check. Note thatrepo_checkhas an argument of the same name that does something related but distinct: there it isTRUEby default, and it lists each archive’s contents in place of the archive itself, without downloading anything. -
max_file_size/max_download_size— size caps (in MB) that gate what may be fetched, so one runaway file or one enormous repository cannot fill your disk. These are an upfront, all-or-nothing gate per repository: if a file or a repository’s total exceeds the cap, that repository is refused (nothing downloaded) and the reason is reported.
These arguments decide what enters the cache; once a file is in, it stays and is reused. See the data_check chapter for full details and examples.
6.3.5 Inspecting and clearing it
Three functions report on the cache and one clears it:
repo_cache_dir() # where it is
repo_cache_size() # how large it is, in bytes
metacheck_cache_info() # location and size of both caches at once
repo_cache_clear() # delete everything in itrepo_cache_clear() also takes a single repository URL, so you can discard one repository’s files while keeping the rest — useful when a repository has changed on the server and you want its files fetched afresh:
repo_cache_clear("https://osf.io/abc123")Because entries are keyed by repository URL and relative path, clearing is always safe: anything still needed is simply downloaded again on the next run.
6.4 How the caches work together in a full run
When you run several modules on a paper — or build an archive across a corpus — the two caches complement each other:
-
repo_checklists the repository’s files (no download). -
data_check/code_checkdownload the files they need; a second module reading the same file in the same run reuses that copy. Withcache = TRUEthe copies are kept for later sessions as well. - When an LLM is enabled,
data_check,codebook_check, and other LLM-using modules send text tollm(), whose answers are stored in the LLM response cache. - Re-running the whole pipeline replays the LLM cache, so no answer is paid for twice. It replays the downloads too if you ran with
cache = TRUE; without it, the files are fetched again, because the previous session’s copies were discarded. This is the main reason to turn the download cache on for an archive build you expect to resume.
Modules such as codebook_check do not have a separate results cache of their own — their re-run savings come entirely from these two shared caches (the files they read and the LLM answers they request). This keeps the caching story simple: there are two on-disk caches, and every module benefits from them without needing its own.
6.5 A note on the in-memory synonym cache
One module, funding_check, builds a list of organisation-name synonyms once and keeps it in memory for the rest of the session (an ordinary in-memory cache, not on disk). It is an implementation detail with no user-facing controls: it simply avoids rebuilding the same lookup repeatedly within a single R session, and disappears when the session ends. You never need to manage it.
6.6 Summary
| Cache | Scope | On by default | Key | Toggle | Clear |
|---|---|---|---|---|---|
| LLM responses | on disk, cross-session | yes | model + prompt + text + type + params | llm_cache() |
llm_cache_clear() |
| Repository files | on disk, cross-session |
no (cache = TRUE to enable) |
repo URL + relative path |
cache argument of data_check / code_check
|
repo_cache_clear() |
| Funding synonyms | in memory, per session | yes | — | — | ends with the session |
The guiding principle is the same for both on-disk caches: cache only what is reproducible, and key it on everything that could change the result. LLM answers are cached because temperature 0 makes them reproducible; downloaded files are cached because a byte-for-byte copy of a remote file is trivially reproducible. In both cases, changing the thing that matters — a prompt, a model, a repository — produces a fresh result automatically, so the caches speed you up without ever giving you a stale answer.
