6  Caching and Reuse

The two slowest and most expensive parts of a Metacheck run are querying a language model and downloading files from an online repository. Both are re-encountered constantly: you re-run a report after tweaking a module, you process the same paper in several modules, or you build an archive from hundreds of repositories overnight and it stops halfway. To avoid paying for the same work twice, Metacheck keeps two on-disk caches — one for LLM responses, one for downloaded repository files — plus a small in-memory cache used by a single module.

This chapter explains what each cache stores, where it lives, how entries are keyed (i.e. what counts as “the same” request), when it is reused, and how to inspect or clear it.

TipThe short version
  • LLM responses are cached automatically (on by default). Re-running a report on the same paper replays the stored answers instead of re-billing the provider. Toggle with llm_cache(); wipe with llm_cache_clear().
  • Downloaded repository files are not cached by default. data_check and code_check download into a temporary folder that is discarded when the R session ends. Pass cache = TRUE to keep the files instead, and clear them with repo_cache_clear().
  • Both caches live in the current working directory, so a project’s cache sits with the project. See where they are with metacheck_cache_info().

6.1 Where the caches live

Both on-disk caches default to a folder in your current working directory, rather than a hidden per-user location elsewhere on the machine. This is deliberate: these caches are per-analysis rather than shared between projects, so keeping them beside the project makes them easy to find, easy to inspect, and easy to delete, and they disappear when you delete the project folder. A hidden cache is exactly the kind that gets silently lost.

Cache Folder (in the working directory) Contents
LLM responses .metacheck_llm_cache/ one .rds per distinct LLM call
Repository files .metacheck_repo_cache/ one sub-folder per repository, mirroring its files

You never need to create these directories by hand; Metacheck creates them on first use. To see where they are and how much space they take, use metacheck_cache_info(), which returns the location and size of both:

metacheck_cache_info()

Because the folder names begin with a dot, they are hidden in most file browsers, and because they sit inside your project you may want to keep them out of version control. Metacheck does not write to your .gitignore for you, so add a line yourself if you use git:

.metacheck_*

You can move both caches somewhere else by setting one option, which is useful if you want a single shared cache across several projects:

options(metacheck.cache.dir = "~/metacheck_caches")

Each cache also has its own override for finer control: the metacheck.repo_cache.dir option for the repository files, and the METACHECK_LLM_CACHE_DIR environment variable for the LLM responses.

6.2 The LLM response cache

6.2.1 Why it is safe to cache an LLM answer

Normally you would not cache a language-model call: ask the same question twice and you can get two different answers. Metacheck sidesteps this by forcing temperature 0 on every llm() call. At temperature 0 the model is (as close as providers allow) deterministic: the same inputs produce the same output. That makes the answer reproducible, and a reproducible answer is one you can safely store and replay. The cache is therefore not a shortcut that changes results — replaying a cached answer gives you the same thing a fresh call would have.

This is the same property that makes Metacheck’s LLM use reproducible in the first place: it is documented alongside the LLM workflow in the Using Large Language Models chapter.

6.2.2 What is stored, and how entries are keyed

The cache is keyed by everything that determines the answer. When llm() is about to make a call, it builds a key from five inputs:

  • the model name,
  • the system prompt,
  • the input text (the row of text being classified),
  • the type specification (the structure the answer must take), and
  • any params (e.g. a seed).

These are serialised to a canonical byte stream and reduced to an MD5 hash, which becomes the filename (<hash>.rds). Two details make the key robust:

  • params are sorted by name before hashing, so list(seed = 1, top_p = 1) and list(top_p = 1, seed = 1) hit the same entry — list order does not matter.
  • The type spec is reduced to its printed form, so if you change the requested output structure (add a field, change a description), the key changes and you get a fresh call rather than a stale answer for the old structure.

Change any of these five inputs and the key changes, so the cache misses and a real call is made. This is exactly what you want: a different question is a different question.

Each entry stores not just the parsed answer (the data frame llm() returns) but also the raw provider result and, where the provider exposes it, the model’s reasoning/thinking trace — so nothing is lost by going through the cache.

6.2.3 Turning it on and off

The cache is on by default. Query or set it with llm_cache():

llm_cache()          # is the cache on? (TRUE by default)
llm_cache(FALSE)     # force every call to hit the provider afresh
llm_cache(TRUE)      # re-enable

Disabling it is mostly useful when you are deliberately testing provider behaviour or measuring latency, and want to be sure no answer is being replayed.

6.2.4 Clearing it

To wipe every stored answer — for example after upgrading a model, or to reclaim disk space — use llm_cache_clear(), which deletes all cache entries and returns how many it removed:

llm_cache_clear()    # returns the number of entries deleted

You do not normally need to clear the cache to pick up a changed prompt or model: because those are part of the key, a change produces a fresh call automatically. Clearing is for reclaiming space or forcing a clean slate.

6.2.5 What this means in practice

Re-running a report on the same paper is close to free on the LLM side: the first run populates the cache, and every subsequent run replays it. This is what makes it comfortable to iterate on a module’s non-LLM code, or to regenerate a report, without watching your API bill climb. It also means the llm_max_calls() cap counts actual provider calls — cached replays do not consume it.

6.3 The repository-file download cache

6.3.1 What it does

Modules that read file contents — data_check and code_check — need local copies of the files a repository shares. repo_check only lists the files on OSF, GitHub, Zenodo and PsychArchives (it does not fetch them). When a content-reading module needs a file, it is downloaded once and read from there, and every later read within the same run reuses that copy instead of downloading again.

Whether those copies survive the session is up to you, and by default they do not. With cache = TRUE, the files are kept in the cache folder and reused by later sessions as well, which is what makes a long archive-building run cheap to recover from after an interruption: the files already fetched are still on disk, and resuming skips straight past them.

6.3.2 How entries are keyed

Each repository gets its own sub-folder, named by a filesystem-safe encoding of its URL (the scheme is stripped and any awkward characters are replaced), so two different repositories never collide and the same repository always resolves to the same folder across sessions. Inside that folder, each file is stored at its repository-relative path, not just its basename — so two files both called data.csv in different sub-folders are kept apart rather than overwriting each other.

Reuse is decided by a simple, robust test: if the target path already exists and is non-empty, it is treated as already downloaded and reused. A zero-byte file (the fingerprint of a download that failed partway) is not trusted — it is re-fetched. This means a broken partial download never silently poisons later runs.

6.3.3 Turning it on

Unlike the LLM cache, which is on by default, the download cache is off by default. Both data_check and code_check take a cache argument that is FALSE unless you say otherwise:

# the default: files go to a temporary folder, discarded when R exits
module_run(paper, "data_check")

# keep the files, and reuse them in later sessions
module_run(paper, "data_check", cache = TRUE)

With cache = FALSE the files are still downloaded and still read; they are simply written to a session folder that R removes on exit, so nothing accumulates on your disk. Files fetched earlier in the same session are reused either way.

Turn it on when you are checking the same repositories repeatedly, or building an archive over several runs, and the download is worth keeping. Leave it off for a one-off check of somebody else’s repository, where you would otherwise be left with copies of their files you never asked to keep.

Reuse of a cached file is safe because a downloaded file is a byte-for-byte copy of the remote one. What you can control separately is how much gets downloaded in the first place, through the arguments described below.

6.3.4 Controlling what gets downloaded

The download cache stores whatever the modules fetch, and data_check gives you fine control over that:

  • download — "data" (the default) fetches only the machine-readable files the checks analyse (tabular data plus codebook/README files); "all" fetches every file in the repository, which is the right choice when building a complete archive; FALSE/"none" downloads nothing and classifies files by name only.
  • skip_types — a vector of file types never to download even under download = "all". When left as NULL, the default, "materials" and "unknown" are skipped: stimuli and software, and the classifier’s collection of manuscripts, logs and unrecognised content, which would only inflate an archive. Pass an explicit vector to override this, or character(0) to download every type. The six types are "data", "code", "documentation", "materials", "output" and "unknown". Skipped files still appear in the manifest, with the reason.
  • peek_zips — look inside each .zip using an HTTP range request, without downloading the whole archive, and only fetch zips that actually contain data or a codebook, leaving a zip of pure stimuli to be linked instead. It is FALSE by default in data_check. Note that repo_check has an argument of the same name that does something related but distinct: there it is TRUE by default, and it lists each archive’s contents in place of the archive itself, without downloading anything.
  • max_file_size / max_download_size — size caps (in MB) that gate what may be fetched, so one runaway file or one enormous repository cannot fill your disk. These are an upfront, all-or-nothing gate per repository: if a file or a repository’s total exceeds the cap, that repository is refused (nothing downloaded) and the reason is reported.

These arguments decide what enters the cache; once a file is in, it stays and is reused. See the data_check chapter for full details and examples.

6.3.5 Inspecting and clearing it

Three functions report on the cache and one clears it:

repo_cache_dir()     # where it is
repo_cache_size()    # how large it is, in bytes
metacheck_cache_info()  # location and size of both caches at once

repo_cache_clear()   # delete everything in it

repo_cache_clear() also takes a single repository URL, so you can discard one repository’s files while keeping the rest — useful when a repository has changed on the server and you want its files fetched afresh:

repo_cache_clear("https://osf.io/abc123")

Because entries are keyed by repository URL and relative path, clearing is always safe: anything still needed is simply downloaded again on the next run.

6.4 How the caches work together in a full run

When you run several modules on a paper — or build an archive across a corpus — the two caches complement each other:

  1. repo_check lists the repository’s files (no download).
  2. data_check / code_check download the files they need; a second module reading the same file in the same run reuses that copy. With cache = TRUE the copies are kept for later sessions as well.
  3. When an LLM is enabled, data_check, codebook_check, and other LLM-using modules send text to llm(), whose answers are stored in the LLM response cache.
  4. Re-running the whole pipeline replays the LLM cache, so no answer is paid for twice. It replays the downloads too if you ran with cache = TRUE; without it, the files are fetched again, because the previous session’s copies were discarded. This is the main reason to turn the download cache on for an archive build you expect to resume.

Modules such as codebook_check do not have a separate results cache of their own — their re-run savings come entirely from these two shared caches (the files they read and the LLM answers they request). This keeps the caching story simple: there are two on-disk caches, and every module benefits from them without needing its own.

6.5 A note on the in-memory synonym cache

One module, funding_check, builds a list of organisation-name synonyms once and keeps it in memory for the rest of the session (an ordinary in-memory cache, not on disk). It is an implementation detail with no user-facing controls: it simply avoids rebuilding the same lookup repeatedly within a single R session, and disappears when the session ends. You never need to manage it.

6.6 Summary

Cache Scope On by default Key Toggle Clear
LLM responses on disk, cross-session yes model + prompt + text + type + params llm_cache() llm_cache_clear()
Repository files on disk, cross-session no (cache = TRUE to enable) repo URL + relative path cache argument of data_check / code_check repo_cache_clear()
Funding synonyms in memory, per session yes — — ends with the session

The guiding principle is the same for both on-disk caches: cache only what is reproducible, and key it on everything that could change the result. LLM answers are cached because temperature 0 makes them reproducible; downloaded files are cached because a byte-for-byte copy of a remote file is trivially reproducible. In both cases, changing the thing that matters — a prompt, a model, a repository — produces a fresh result automatically, so the caches speed you up without ever giving you a stale answer.