Inside Clean Room: A Reproducibility Boundary in Python, R, and WebAssembly

Clean Room is GoFigr’s answer to a specific question: what would it take for a figure to be re-runnable by someone who isn’t you — different machine, no Python or R installed, months later, with different parameters?

The obvious answer is “ship the notebook,” and it’s wrong. A notebook is a record of exploration, and exploration accumulates state: variables defined forty cells ago, a helper function you tweaked and re-ran, a DataFrame that only exists because you loaded it interactively last Tuesday. A figure produced in that environment depends on all of it, and none of it is written down anywhere.

Clean Room’s design premise is that the unit of reproducibility is not the notebook — it’s one function. If we can prove, mechanically, that a function depends on nothing but its arguments and its declared packages, then we can serialize those arguments, record those packages, and re-run the function anywhere. Including inside a WebAssembly interpreter in a stakeholder’s browser.

This post is about the machinery: how we enforce that boundary in Python, how we enforce it in R (where the obvious mechanism doesn’t exist), and what it takes to actually re-execute the result in a browser with no server-side compute.

The boundary, precisely

When you write this in a notebook:

@reproducible(interactive=True)
def flipper_hist(data, bins: int = SliderParam(20, min=5, max=100)):
sns.histplot(data, x="flipper_length_mm", bins=bins)
flipper_hist(penguins)

three things cross the boundary, and nothing else:

  1. Declared parameters — serialized and stored with the revision
  2. Passed data — DataFrames, serialized as Parquet
  3. Declared packages — recorded with exact installed versions

Notebook globals, local files, and hidden state stay behind. Critically, this is enforced at runtime, not by convention: if flipper_hist references a global, it doesn’t produce a subtly-irreproducible artifact — it raises NameError on the first run, on your machine, while you can still fix it.

One thing this is not is a security sandbox. Inside the function you can import os in Python or reach any namespace via :: in R. The boundary exists to catch accidents — and accidental dependence on ambient state is essentially the entire reproducibility failure mode. Malicious code is out of scope; you wrote the function.

Python: rebind the code object

CPython gives us a remarkably clean mechanism for this. A function object is a small wrapper around a compiled code object plus a reference to its globals dict. types.FunctionType lets you reassemble one with different globals:

def _run_clean(func, params, packages, ...):
clean_globals = {"__builtins__": __builtins__}
for alias, module_name in packages.items():
clean_globals[alias] = importlib.import_module(module_name)
isolated_func = types.FunctionType(
func.__code__, # same compiled bytecode
clean_globals, # builtins + declared packages, nothing else
name=func.__name__,
argdefs=func.__defaults__,
closure=func.__closure__,
)
return isolated_func(**params)

The function’s bytecode is untouched — same LOAD_GLOBAL instructions — but every global lookup now resolves against a dict containing only builtins and the packages you declared ({"pd": pandas, "np": numpy, "plt": matplotlib.pyplot, "sns": seaborn} by default, extensible via the packages= argument). Reference anything else and LOAD_GLOBAL fails with NameError.

There’s one honest caveat visible in that snippet: closure=func.__closure__. If you define a @reproducible function nested inside another function, its closure cells survive the rebinding — closures are carried by the function object, not the globals dict. In practice @reproducible functions are defined at notebook top level where the closure is empty, but the boundary is a globals boundary, not a total-environment one.

Parameters are widget declarations

The decorator inspects the signature with inspect.signature(...).bind(...) and treats default values as widget descriptors. SliderParam(20, min=5, max=100) is a default value that is its own UI spec; the decorator unwraps it to 20 before calling the function and keeps the descriptor for the manifest. Plain defaults get inferred: bool → checkbox, int/float → slider, str → text box, and a Literal["yes", "no", "auto"] type hint becomes a dropdown via typing.get_args.

Run what you stored, not what you have

The subtlest design decision in the Python client is a round trip that looks pointless:

def round_trip_params(params):
bundle = serialize_params(params) # DataFrames -> Parquet bytes
return deserialize_params(bundle.manifest, bundle.dataframes)

Before the function runs — even the very first time, on your machine — every parameter is serialized to its storage format and immediately deserialized back, and the function runs on the round-tripped values. Your DataFrame goes through Parquet before your own eyes ever see the plot.

The reason: Parquet is not a perfect image of a pandas DataFrame. If a dtype doesn’t survive, or a categorical collapses, or an index behaves differently, we want that visible in the notebook on run one — not discovered in a browser re-run six months later, where it would look like Clean Room corrupted the figure. If the round trip changes your output, that’s the system telling you your artifact wasn’t storable as-is.

The distribution-name problem

The manifest records package versions, and here Python has a trap: the name you import is not the name you install. import sklearn comes from the scikit-learn distribution; import cv2 from opencv-python. The manifest must record distribution names, because that’s what the browser’s package installer will consume later:

from importlib.metadata import packages_distributions
dist_map = packages_distributions() # {"sklearn": ["scikit-learn"], ...}

Get this wrong and the browser runtime would try micropip.install("sklearn") and fail. The mapping happens once, at capture time, on the machine where the truth is known.

R: there are no decorators here

R has no decorator syntax, and more fundamentally, an R closure carries its enclosing environment with it — you can’t “swap the globals” of a function the way types.FunctionType does, because name resolution walks the environment chain the function was born with.

So the R implementation (gofigR) inverts the pattern: reproducible() is not a wrapper that returns a modified function. It’s a runner that takes your function apart and evaluates its body somewhere else:

reproducible(
function(data = static(penguins),
bins = slider(20L, min = 5L, max = 100L),
species = dropdown("Adelie", choices = c("Adelie", "Chinstrap", "Gentoo"))) {
p <- ggplot(data[data$species == species, ], aes(flipper_length_mm)) +
geom_histogram(bins = bins)
publish(p, figure_name = "Flipper Histogram")
},
packages = c("ggplot2")
)

Inside, the boundary is an evaluation environment whose parent skips globalenv() entirely:

clean_env <- new.env(parent = baseenv()) # chain: clean_env -> base R. No globals.
for (pkg in packages)
for (nm in getNamespaceExports(pkg))
assign(nm, get(nm, envir = asNamespace(pkg)), envir = clean_env)
clean_env$publish <- gofigR::publish
# ...inject resolved parameter values...
eval(body(fn), envir = clean_env)

Because the parent is baseenv() rather than globalenv(), lookups from inside the body resolve against the clean environment, then base R — your workspace variables simply aren’t on the search path. The function’s own enclosing environment is bypassed too, since we evaluate body(fn) directly rather than calling the closure.

Two R-specific mechanics are worth savoring:

Parameters live in the formals. R keeps default arguments as unevaluated expressions, so the widget declarations sit in formals(fn) and are harvested by evaluating each default and checking for the gf_param class. Same idea as Python’s SliderParam-as-default, but implemented through R’s lazy-default machinery. Non-widget defaults are auto-wrapped as static parameters.

publish() finds the clean room through the environment chain. The runner plants a sentinel in the clean environment (clean_env$.gf_clean_room_context <- context), and publish() — which knows nothing about reproducible() as a matter of API — discovers it at call time:

ctx <- get0(".gf_clean_room_context", envir = parent.frame(), inherits = TRUE)

inherits = TRUE walks the chain from the call site; if publish() was called inside a clean room, the sentinel is an ancestor and the revision gets the source, manifest, and Parquet data attached. Called outside, get0 returns NULL and publish() behaves normally. It’s the R idiom for what Python does with contextvars — context passed through dynamic extent rather than arguments.

Source capture mirrors Python’s approach in spirit: where Python uses ast to strip the decorator and def line and keep the body, R uses body(fn) — already an AST — and deparses its statements, skipping the {. No regex in either language.

The manifest: one contract, two producers, one consumer

Both clients emit the same JSON manifest:

{
"language": "python",
"language_version": "3.12.4",
"function_name": "flipper_hist",
"packages": {"seaborn": "0.13.2", "pandas": "2.2.3"},
"imports": {"sns": "seaborn", "pd": "pandas"},
"parameters": {
"bins": {"type": "integer", "widget": "slider", "value": 20, "min": 5, "max": 100, "step": 1},
"species": {"type": "string", "widget": "dropdown", "value": "Adelie", "choices": ["Adelie", "Chinstrap", "Gentoo"]},
"data": {"type": "dataframe"}
},
"rendering": {"dpi": 100}
}

Parameter types are language-neutral (integer, number, boolean, string, none, dataframe) — no int64 vs numeric. DataFrames carry no value in the manifest; their bytes are stored as separate Parquet objects on the revision, referenced by parameter name. A revision flagged is_clean_room carries exactly three kinds of extra data: the function source, this manifest, and zero or more Parquet tables. That’s the entire contract the browser needs.

(The rendering block exists because a re-run should look like the original: the Python client records the session’s matplotlib DPI so the browser can match it.)

The browser: re-running with no server

When someone opens a Clean Room figure, the studio boots a real language runtime in a Web Worker — Pyodide for Python, webR for R — and replays the manifest. Nothing executes server-side; the WASM interpreter, the packages, and your data all live in the visitor’s browser.

The choreography, in order:

Packages install on demand. Pyodide uses micropip with the distribution names recorded in the manifest — this is where the sklearn/scikit-learn mapping pays off. On top of the user’s packages, the studio installs its own infrastructure: pyarrow (to read the Parquet parameters) and matplotlib (to capture output). webR installs from the webR binary repo, with one notable substitution: Arrow doesn’t build for WASM, so the R side reads Parquet with nanoparquet — the same reason the R client writes with nanoparquet at capture time. Import aliases from the manifest (sns → seaborn) are recreated after install.

Data loads from the revision, not from you. Each Parquet object is fetched from the API and materialized in the runtime — pd.read_parquet(io.BytesIO(...)) on the Python side; on the R side the bytes are written into webR’s virtual filesystem and read back with nanoparquet::read_parquet(), then passed through as.data.frame() to strip the tibble class (webR’s vctrs/pillar builds and ggplot2 disagree otherwise).

Each run gets a fresh namespace, not a fresh interpreter. Booting Pyodide and installing packages is the expensive part, so the worker does it once, then snapshots the post-setup namespace (dict(globals()); as.list(globalenv()) in R). Every parameter tweak re-executes against a clone of that snapshot, so runs can’t leak state into each other — the same isolation philosophy as the original boundary, applied per-run.

return needs a home. The stored source is a function body, and bodies may contain top-level return. The Python worker wraps the code in a synthetic def __cleanroom_main(): ... and calls it, so the statement stays legal.

Output capture is a monkeypatch. In Pyodide, matplotlib is forced to the Agg backend and plt.show is replaced with a capture routine that walks plt.get_fignums(), renders PNGs at the manifest’s DPI, and pads the image to leave room for the QR watermark. In R there’s no display to intercept, so the worker injects a publish() shim that dispatches on plot class — ggsave for ggplot objects, a PNG device plus replayPlot for recorded plots. That shim guards against our favorite R footgun in the codebase: base-graphics arguments arrive as unevaluated promises, and the shim must attempt grid conversion before calling inherits() on the argument — because inherits() would force the promise and draw the plot on the wrong device.

Saving is a derivation, not an overwrite. When a visitor publishes their re-run, the studio calls a derive endpoint on the source revision. The server clones the immutable inputs (notably the Parquet data), and the studio appends the new images, the current code, and a manifest updated with the current parameter values. The result is a new revision linked to its parent — provenance is a chain of derivations, each carrying the exact parameters that produced it, watermarked with a QR pointing at its own ID. Nobody’s original is ever touched.

Degrade, never break

A capture system that can make your analysis fail is a capture system you’ll turn off. So every unsupported case in both clients resolves the same way — warn, skip the clean room, run the function anyway:

Situation Behavior
Parameter isn’t serializable (model object, lambda, environment) warn, run directly without clean room
DataFrames exceed 100 MB (memory_usage(deep=True) estimate) warn, run directly without clean room
nanoparquet missing (R) warn, run directly without clean room
interactive=True outside Jupyter / inside knitr warn, run once non-interactively
anywidget frontend missing in JupyterLab banner, render figure non-interactively

The invariant: @reproducible never stands between you and your plot. Isolation and capture apply when the preconditions hold; when they don’t, you get your figure plus a warning explaining what to fix.

Summary

The trick that makes Clean Room work is that “reproducible” is checkable. Python checks it by rebinding a code object to a globals dict you control; R checks it by evaluating a function body in an environment whose parent skips the global workspace. Both funnel into one language-neutral contract — source, manifest, Parquet — which is exactly enough for a WASM runtime in a browser to replay the analysis, parameters and all, with no server and no setup.

If you want to poke at the result: here’s a live Clean Room figure you can re-run right now, and the Python and R docs cover the API surface.