When a rerun differs from a reference result, locate the first disagreement. Comparing only the final score mixes input changes, implementation differences and presentation choices into one number.
Four checkpoints
| Checkpoint | Compare | Preserve |
|---|---|---|
| Inputs | File identity, row IDs and relevant timestamps | Input version or checksum |
| Selection | Training and evaluation membership | Exact row IDs, not only counts |
| Computation | Predictions, outcomes and metric definition | Unrounded intermediate results |
| Reporting | Units, rounding and displayed subset | The report-generation rule |
These are diagnostic checkpoints, not a promise that every mismatch has a simple cause. Do not edit the reference output to make a comparison pass.
An original calculation
The reference uses two observed values, 1 and 3, with predictions 0 and 1. Squared errors are 1 and 4, so mean squared error is (1 + 4) / 2 = 2.5.
Your displayed value is approximately 1.581. If the underlying rows and predictions match, check the metric: the square root of 2.5 is approximately 1.581. You may be reporting root mean squared error rather than mean squared error. That is a definition mismatch, not necessarily a different fitted model.
Now change the second prediction to 2. The squared errors become 1 and 1, and mean squared error becomes 1. A matching metric label no longer resolves the difference; inspect how that prediction was produced. Keep the two diagnoses separate in your notes.
A seed is only part of the record
The NumPy compatibility policy places conditions around matching random streams. For a NumPy-based study, preserve the generator choice and relevant environment and call sequence, not just an integer seed. Other libraries have their own guarantees.
For deterministic calculations, also distinguish exact equality from a justified numerical tolerance. State a tolerance before using it to judge agreement and explain why it is suitable for the result. A tolerance that swallows a material error is not useful reproduction evidence.
Write a short mismatch note
Use this structure: expected artifact, observed artifact, earliest disagreement, explanation checked, change made and remaining limitation. For the first example: “Inputs and predictions matched. The reference reported MSE; my display reported RMSE. Relabeling the metric resolved the apparent discrepancy without refitting.”
If the cause remains unknown, say so. A small final-score difference does not prove equivalence, and a large difference does not by itself identify the cause.
Use the guided-lab versus independent-study comparison when choosing your next practice resource. The README guide helps package the resulting instructions. Label supplied code and your own additions separately; recreating a teaching result is not an independent trading discovery.
Read before choosing
Open the actual pages.
7 sample pages, including complete explanations. No email address or account required.
Open the PDF previewPreview page 4 of 7. Use Enlarge page for a closer view. When the page is focused, use left and right arrows to change pages.
