repro lensOPEN SOURCE

FOR ML DEVELOPERS & CODING AGENTS

Change the code.
Keep the
evidence.

Your refactor runs. But did it change the result? Catch reproducibility risks, replay experiments, and compare what actually changed.

Free & open source No API key Runs locally

experiment / change reviewDEMO
ONE EXPERIMENT. THREE EDITS.

A repeatable change can
still change the result.

experiment.py equivalent conditional

− return int(score >= THRESHOLD)+ return 1 if score >= THRESHOLD else 0
Two runs after the edit✓ Matched
Compare with baseline✓ Matched
Fixture accuracy1.00 1.00

Same declared outputs. Code change recorded.

experiment.py classification threshold

− THRESHOLD = 0.5+ THRESHOLD = 0.7
Two runs after the edit✓ Matched
Compare with baseline↗ Mismatch
Fixture accuracy1.00 0.75

Repeatable on its own. Different from the baseline.

pyproject.toml after changing the threshold

− atol = 0.0+ atol = 0.3
Two runs after the edit✓ Matched
Compare with baseline! Not comparable
Fixture accuracy1.00 0.75

The contract changed. Review the new tolerance.

Expected outcomes from our CI-checked demo. Eight synthetic samples; scripted edits. This preview does not execute code.

A successful run is only the beginning.

SELECTED API CHECKS FOR

scikit-learnXGBoostLightGBMPyTorchTensorFlowLightning
See coverage

01 EXPLORE THE CHECKER

Small change.
Important difference.

Pick a case. See what the checker finds.
Take an explicit alternative back to your code.

CHECK EXPLORER16 cases 8 libraries & frameworks

scikit-learn

scikit-learn data split

R101 · warning

Make a shuffled split reproducible.

Before

from sklearn.model_selection import train_test_split
train_test_split(X, y)

R101 · warning · line 2: sklearn.model_selection.train_test_split has no explicit non-None random_state.

One explicit alternative

from sklearn.model_selection import train_test_split
train_test_split(X, y, random_state=seed)

No covered finding in this snippet. This is not proof of repeatability.

The covered split omits random_state. An explicit seed makes the intended policy visible; global RNG initialization can also matter and should be reviewed.

Link to this example
NumPy

NumPy generator

R102 · warning

Avoid an implicitly initialized generator.

Before

import numpy as np
np.random.default_rng()

R102 · warning · line 2: numpy.random.default_rng has no explicit non-None seed.

One explicit alternative

import numpy as np
np.random.default_rng(seed)

No covered finding in this snippet. This is not proof of repeatability.

Calling default_rng without a seed draws fresh entropy. Use a seed chosen for the experiment, and retain its inputs and environment as well.

Link to this example
Python

Python random generator

R103 · warning

Make a local random generator explicit.

Before

import random
random.Random()

R103 · warning · line 2: random.Random has no explicit non-None seed.

One explicit alternative

import random
random.Random(seed)

No covered finding in this snippet. This is not proof of repeatability.

A new Random instance without an argument is not controlled by seeding the module-level generator. Supply the seed intended for this instance.

Link to this example
NumPy

NumPy global shuffle

R111 · review

Shuffle indices from a seeded generator instead of unseeded global state.

Before

import numpy as np
indices = np.arange(len(X))
np.random.shuffle(indices)

R111 · review · line 3: numpy.random.shuffle draws from NumPy's global RNG, which this file never seeds.

One explicit alternative

import numpy as np
rng = np.random.default_rng(seed)
indices = np.arange(len(X))
rng.shuffle(indices)

No covered finding in this snippet. This is not proof of repeatability.

np.random.shuffle draws from NumPy's global RNG, and nothing in this file seeds it. A Generator created with the experiment's seed keeps the policy local and visible. Seeding may also happen in another module; this is a review item, not a defect finding.

Link to this example
NumPy

NumPy seed without a value

R111 · review

Does seed() without a value pin the global RNG?

Before

import numpy as np
np.random.seed()
indices = np.arange(len(X))
np.random.shuffle(indices)

R111 · review · line 4: numpy.random.shuffle draws from NumPy's global RNG, which this file never seeds.

One explicit alternative

import numpy as np
np.random.seed(seed)
indices = np.arange(len(X))
np.random.shuffle(indices)

No covered finding in this snippet. This is not proof of repeatability.

np.random.seed() and seed(None) draw OS or clock entropy, so they do not make a later shuffle repeatable. Pass the experiment's seed, or switch to a Generator created with that seed.

Link to this example
Python

Python global shuffle

R112 · review

Make module-level random state explicit.

Before

import random
random.shuffle(items)

R112 · review · line 2: random.shuffle draws from Python's global random module, which this file never seeds.

One explicit alternative

import random
rng = random.Random(seed)
rng.shuffle(items)

No covered finding in this snippet. This is not proof of repeatability.

random.shuffle draws from Python's module-level generator, and this file never seeds it. A dedicated Random instance keeps the experiment's policy visible; seeding may also happen elsewhere, so this remains a review item.

Link to this example
XGBoost

XGBoost linear updater

R104 · warning

A seed does not make every updater deterministic.

Before

import xgboost as xgb
xgb.XGBRegressor(booster='gblinear', random_state=seed)

R104 · warning · line 2: xgboost.XGBRegressor selects gblinear with the nondeterministic shotgun updater.

One explicit alternative

import xgboost as xgb
xgb.XGBRegressor(booster='gblinear', updater='coord_descent', random_state=seed)

No covered finding in this snippet. This is not proof of repeatability.

The default shotgun updater for gblinear is nondeterministic. Coordinate descent is an alternative to review for your experiment.

Link to this example
LightGBM

LightGBM CPU histogram

R105 · review

Make CPU histogram construction explicit.

Before

import lightgbm as lgb
params = {'deterministic': True}
lgb.train(params, train_data)

R105 · review · line 3: lightgbm.train enables determinism without exactly one forced histogram mode.

One explicit alternative

import lightgbm as lgb
params = {'deterministic': True, 'force_col_wise': True, 'seed': seed}
lgb.train(params, train_data)

No covered finding in this snippet. This is not proof of repeatability.

For the covered CPU configuration, deterministic mode also needs an explicit histogram choice. Select row-wise or column-wise construction for your workload; the example uses column-wise.

Link to this example
PyTorch

PyTorch weight initialization

R113 · review

Make random weight initialization repeatable.

Before

import torch
weights = torch.randn(64, 32)

R113 · review · line 2: torch.randn draws from PyTorch's global RNG, which this file never seeds.

One explicit alternative

import torch
torch.manual_seed(seed)
weights = torch.randn(64, 32)

No covered finding in this snippet. This is not proof of repeatability.

torch.randn draws from PyTorch's global RNG, which this file never seeds. torch.manual_seed with the experiment's seed makes the initialization repeatable on one device; cuDNN kernels and data loading still need their own review.

Link to this example
PyTorch

PyTorch shuffled data

R106 · review

Give shuffled loading an explicit generator.

Before

from torch.utils.data import DataLoader
DataLoader(dataset, shuffle=True)

R106 · review · line 2: torch.utils.data.DataLoader samples data without an explicit seeded generator.

One explicit alternative

import torch
from torch.utils.data import DataLoader
DataLoader(dataset, shuffle=True, generator=torch.Generator().manual_seed(seed))

No covered finding in this snippet. This is not proof of repeatability.

A missing generator needs review because RNG state may be managed elsewhere. The example makes generator ownership explicit; worker initialization and dataset behavior still need attention.

Link to this example
PyTorch

PyTorch random sampler

R106 · review

Does a sampler need its own generator?

Before

from torch.utils.data import DataLoader, RandomSampler
DataLoader(dataset, sampler=RandomSampler(dataset))

R106 · review · line 2: torch.utils.data.RandomSampler samples data without an explicit seeded generator.

One explicit alternative

import torch
from torch.utils.data import DataLoader, RandomSampler
DataLoader(dataset, sampler=RandomSampler(dataset, generator=torch.Generator().manual_seed(seed)))

No covered finding in this snippet. This is not proof of repeatability.

DataLoader does not shuffle when a sampler is passed, so an unseeded RandomSampler is invisible to the shuffle check. The sampler draws from the global RNG until generator= is a seeded torch.Generator. SequentialSampler is not random and is not flagged.

Link to this example
PyTorch

PyTorch unseeded generator

R106 · review

Does passing torch.Generator() seed the shuffle?

Before

import torch
from torch.utils.data import DataLoader
DataLoader(dataset, shuffle=True, generator=torch.Generator())

R106 · review · line 3: torch.utils.data.DataLoader samples data without an explicit seeded generator.

One explicit alternative

import torch
from torch.utils.data import DataLoader
DataLoader(dataset, shuffle=True, generator=torch.Generator().manual_seed(seed))

No covered finding in this snippet. This is not proof of repeatability.

DataLoader accepts a generator object, but the constructor draws OS entropy until manual_seed is given a value. Chain manual_seed with the experiment's seed. A generator variable is still accepted without proving that it was seeded.

Link to this example
PyTorch

PyTorch cuDNN benchmark

R107 · warning

Review cuDNN algorithm benchmarking.

Before

import torch
torch.backends.cudnn.benchmark = True

R107 · warning · line 2: cuDNN benchmarking can select different algorithms across runs.

One explicit alternative

import torch
torch.backends.cudnn.benchmark = False

No covered finding in this snippet. This is not proof of repeatability.

Benchmarking can select different convolution algorithms between runs. Turning it off addresses this setting, but does not establish deterministic execution.

Link to this example
TensorFlow

TensorFlow RNG state

R108 · warning

Choose how the RNG state is initialized.

Before

import tensorflow as tf
tf.random.Generator.from_non_deterministic_state()

R108 · warning · line 2: tensorflow.random.Generator.from_non_deterministic_state explicitly initializes nondeterministic RNG state.

One explicit alternative

import tensorflow as tf
tf.random.Generator.from_seed(seed)

No covered finding in this snippet. This is not proof of repeatability.

This constructor explicitly requests nondeterministic state. The alternative uses your chosen seed; it does not configure every TensorFlow operation.

Link to this example
Lightning

Lightning Trainer

R109 · review

Decide whether determinism is required or advisory.

Before

from lightning.pytorch import Trainer
Trainer(deterministic='warn')

R109 · review · line 2: lightning.pytorch.Trainer does not request strict determinism with benchmarking disabled.

One explicit alternative

from lightning.pytorch import Trainer, seed_everything
seed_everything(seed, workers=True)
Trainer(deterministic=True, benchmark=False)

No covered finding in this snippet. This is not proof of repeatability.

Warning-only mode allows operations without a deterministic implementation to proceed. Strict mode can raise errors, so review the choice and its effect on training.

Link to this example
PyTorch

PyTorch algorithm policy

R110 · review

Make unsupported deterministic operations visible.

Before

import os
import torch
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
torch.use_deterministic_algorithms(True, warn_only=True)

R110 · review · line 4: torch.use_deterministic_algorithms allows operations without a deterministic implementation.

One explicit alternative

import os
import torch
os.environ["CUBLAS_WORKSPACE_CONFIG"] = ":4096:8"
torch.use_deterministic_algorithms(True, warn_only=False)

No covered finding in this snippet. This is not proof of repeatability.

With warn_only=True, an unsupported operation can continue. Strict mode changes that behavior and may stop a run; it is a policy decision, not a universal fix.

Link to this example
CHECKED AT BUILD TIME

These are saved static results from the same checker as the CLI. To scan your own code, use the live check below. seed is a value you choose for your experiment.

Useful for your next experiment? Keep Repro Lens in your GitHub stars so you can find it again.

Star on GitHub

02 CHECK YOUR CODE

Paste your script.
See what it leaves to chance.

The same checker as the CLI, running in this tab.
Your code is not uploaded anywhere.

LIVE CHECKPython in your browser nothing uploaded

Press Check code. The first check downloads Python for this page (about 13 MB, once).

LIVE, IN YOUR BROWSER

Runs the analyzer from this build with Pyodide, served from this site. It screens selected APIs as text and never runs your code. A clean result is not proof of repeatability.

03 FROM CODE TO EVIDENCE

Three questions.
One review workflow.

Use it with your coding agent
01 / BEFORE THE RUN

Check the code.

Spot known reproducibility risks with file locations. No ML dependencies needed.

repro-lens checkReview your first findings
02 / RUN IT AGAIN

Replay the experiment.

Run your configured command twice. Retain metrics, artifact hashes and logs.

repro-lens verifyDefine what should match
03 / AFTER THE EDIT

Compare the evidence.

See whether a change preserves baseline outputs or needs a closer review.

repro-lens compare before.json after.jsonTry the scikit-learn CPU example

A match applies to the declared outputs and recorded environment. It does not establish scientific validity.

04 MAKE IT YOURS

Bring your code.
Leave with evidence.

Start with the scripted demo. Then check a training script or review your agent’s next refactor.

Python 3.11+ · Git · uv
No ML libraries or API key for this demo.

TERMINAL
git clone https://github.com/00200200/repro-lens.git
cd repro-lens
uv run --no-dev python examples/agent_review/demo.py

Run from a directory where repro-lens does not already exist. The demo runs locally and retains its reports.

Install the CLI for your own project

BUILT IN THE OPEN

One useful example
can help the next person.

Found a false alarm, a missed call, or a better example? That’s a good contribution.