What was actually checked
Set the boundary before sending a prompt
- 01
Start from a known runtime
Follow the Alpha quick start for 0.1.3-alpha.1 and its fixed source commit. Select your configured model in Settings → Models. Credentials belong in provider configuration, not prompts or test fixtures.
- 02
Separate the project from the runtime
Keep the exercise in its own directory and select that directory in the Web UI. Do not create the official source runtime inside this website or another parent with conflicting node_modules.
- 03
Keep the first task narrow
Inspect effective sandbox/approval settings. Use the disposable workspace, permit only the work required by the prompt, and avoid adding plugins or dependencies to solve a three-file problem.
Create a small project with a visible failure
Use Node.js 24 and a Bash-compatible terminal in a new disposable directory. These commands work as a Bash exercise on macOS/Linux; on Windows use Git Bash or WSL. The new harness-lab directory must not already exist. There are no npm dependencies, accounts, private data or production files in the fixture. The deliberately simple CSV contains no quoted commas; this exercise is not a general CSV parser.
create_harness_lab() {
mkdir harness-lab || return
cd harness-lab || return
cat > orders.csv <<'CSV'
id,status,cents
A,paid,1200
B,refunded,700
C,paid,800
D,pending,400
CSV
cat > report.mjs <<'JS'
export function totalPaid(csv) {
const rows = csv.trim().split(/\r?\n/).slice(1);
return rows.map(row => row.split(','))
.filter(([, status]) => status !== 'pending')
.reduce((sum, [, , cents]) => sum + Number(cents), 0);
}
JS
cat > report.test.mjs <<'JS'
const { default: assert } = await import('node:assert/strict');
const { readFileSync } = await import('node:fs');
const { default: test } = await import('node:test');
const { totalPaid } = await import('./report.mjs');
const csv = readFileSync(new URL('./orders.csv', import.meta.url), 'utf8');
test('only paid orders count', () => assert.equal(totalPaid(csv), 2000));
test('refunds alone count as zero', () => assert.equal(totalPaid('id,status,cents\nB,refunded,700\n'), 0));
test('an empty ledger counts as zero', () => assert.equal(totalPaid('id,status,cents\n'), 0));
JS
cp report.mjs report.original.mjs
node --test report.test.mjs
}
create_harness_labThe first run should fail two assertions and pass one. The buggy code counts the 700-cent refunded order, producing 2700 instead of 2000. The rule is exact: count only rows whose status is paid; refunded and pending rows contribute zero. All values use integer cents.
Ask for a reproducible diagnosis first
Inspect only orders.csv, report.mjs and report.test.mjs. Do not change files yet. State the accounting rule implied by the tests, identify the refunded row, and run node --test report.test.mjs. Explain why the current filter yields 2700 instead of 2000. If you cannot run the command, report the exact environment problem without inventing output.A useful diagnosis names row B, the status !== 'pending' filter and the two failing assertions. 'The tests should pass after a fix' is not a test result. If the agent proposes a new CSV library or a broad refactor, redirect it to the existing fixture and acceptance criteria.
Allow the focused repair, then check it independently
In this disposable project, fix report.mjs so totalPaid counts only paid rows. Read orders.csv and report.test.mjs first, run node --test report.test.mjs and report the original failure. Do not edit orders.csv, report.test.mjs or report.original.mjs; do not install packages or access other directories. Make the smallest implementation change, rerun all three assertions, and calculate the total from orders.csv. Return the changed file, the observed test result, the total and any limitation. If a command fails for an environmental reason, distinguish that from a failing assertion.node --test report.test.mjs
node --input-type=module -e "import {readFileSync} from 'node:fs'; import {totalPaid} from './report.mjs'; console.log(totalPaid(readFileSync('orders.csv','utf8')))"
diff -u report.original.mjs report.mjsExpect three passing tests and the separate total 2000. The diff should normally change the filter to status === 'paid'. A different implementation is acceptable if it meets the same rule and preserves the tests. diff returns exit code 1 when files differ; that is expected here, not a failed test. A missing Node executable or permission denial is an environment failure, not evidence that the calculation is wrong.
.filter(([, status]) => status === 'paid')The locally checked reference fix changes only the filter. The input and tests remain unchanged. Keep the unit-test output and the separate calculated total together; do not accept a screenshot of a green terminal without the command and the relevant file diff.
Second task: a changed setting keeps the old cached instance
Create these two additional files in the same disposable project. The account stays demo, but mode changes from light to dark. The cache currently compares only account, so it incorrectly returns the old object. This small example teaches a common diagnostic pattern: changing configuration and invalidating a constructed instance are different operations.
create_cache_lab() {
test ! -e cache.mjs && test ! -e cache.test.mjs && test ! -e cache.original.mjs || return
cat > cache.mjs <<'JS'
let cached;
export function getSettings(config) {
if (!cached || cached.account !== config.account) cached = { ...config };
return cached;
}
JS
cat > cache.test.mjs <<'JS'
const { default: assert } = await import('node:assert/strict');
const { default: test } = await import('node:test');
const { getSettings } = await import('./cache.mjs');
test('a changed setting refreshes the instance', () => {
const first = getSettings({ account: 'demo', mode: 'light' });
const second = getSettings({ account: 'demo', mode: 'dark' });
assert.notEqual(second, first);
assert.equal(second.mode, 'dark');
assert.equal(getSettings({ account: 'demo', mode: 'dark' }), second);
});
JS
cp cache.mjs cache.original.mjs
node --test cache.test.mjs
}
create_cache_labInspect cache.mjs and cache.test.mjs. Reproduce the failing test, then fix the cache so changing account OR mode constructs a new instance, while identical values reuse the existing instance. Keep the test unchanged; add a separate assertion for an account change if useful. Do not introduce dependencies, timers or a global cache reset. Run node --test cache.test.mjs and explain which inputs determine identity.if (!cached || cached.account !== config.account || cached.mode !== config.mode) cached = { ...config };Run node --test cache.test.mjs. The checked reference fix passes: the changed mode gives a different object with dark, and the unchanged next call reuses that object. The fixture has only account and mode; in a real service enumerate every construction input, preserve failed-read behavior and verify rotation without exposing secret values. This exercise does not certify a production cache or payment provider.
Why these checks belong in the exercise
Anthropic’s context-engineering article discusses targeted retrieval and structured notes; its evaluation article distinguishes the conversation record from the outcome. Our application is specific: start with the named files, reread saved changes, and keep the real test output and total beside the diff. HANDOFF.md records paths, executed commands and remaining work so a resumed session can check the current files again. This links context, evidence and continuation without treating a longer prompt or log as proof of quality.
Make a paused task easy to resume
Keep report.original.mjs and cache.original.mjs unchanged. Before starting a new session, ask for a short handoff containing the current files, last executed commands and unresolved work. On resumption, read the files and rerun the tests: a session summary may describe an earlier state. DSH’s stored session is history, not a replacement for source-control or file backups.
Write HANDOFF.md for this exercise: objective; files changed; original failures; exact commands actually executed and their results; remaining issue, if any; and the next verification step. Do not include credentials or claim unrun checks. A new session should be able to continue by reading this note and the current files.Recover according to the failure you observed
- 01
A command cannot run
Check the terminal’s Node version, cwd and selected workspace. Correct the environment before changing source to compensate for a missing executable or permission denial.
- 02
The agent changed tests or extra files
Stop the task, preserve its attempted changes and review the diff. Restore only the affected exercise files from their known originals, then restate the acceptance criteria.
- 03
The model or provider is unavailable
Keep the local files and handoff. Correct provider settings outside the prompt, then resume the session or start a new one with the same acceptance criteria. Do not reinterpret an API failure as a successful code test.
cp report.mjs report.attempt.mjs
cp report.original.mjs report.mjs
node --test report.test.mjsAfter restoration, the original two test failures should return. This is an intentional recovery check. Keep report.attempt.mjs for comparison. In your real repository, use a clean branch/worktree and a reviewed file-specific restore; never apply this fixture’s overwrite command to unrelated work.
Turn a successful routine into a small reusable Skill
Only after completing the exercise, write down its inspect → reproduce → edit → test → handoff sequence. DSH’s documented local discovery includes project .dsh/skills and .agents/skills; names and precedence matter. A Skill records a workflow, not new system permissions. Review the file before making it discoverable, keep commands scoped, and do not load an unknown extension merely because it promises automation.
Continue with a concrete next step
Read the primary documentation
Common questions
Before you make a change
Answers about formats, compatibility, evidence, and rollback.
What does this guide verify?
This guide covers: Pinned DSH Alpha sources and primary context-management and evaluation references; Two synthetic Node fixtures, reference repairs and scoped file recovery executed locally; No comparative model benchmark or real provider task result is claimed
What should I do first?
Use two small jobs to establish a repeatable habit: correct an order total, then fix a settings cache that ignores a change. Each job has an input, an observable failure, a copyable prompt and an independent check. Start with the disposable project below before applying the process to valuable work.
Which boundary matters most?
Both Node examples below were run locally, including the original failures, reference repairs and restoration of the original report.mjs module. DSH instructions follow 0.1.3-alpha.1 at d347e703908d0406b7a7ef80e3a0e594d86b2215. These are exercises for your configured model; no model-run result is claimed.
Should I ask the agent to plan every tiny change?
For these fixtures, a short diagnosis and an explicit test command are enough. A larger task benefits from milestones when each milestone has a reviewable output and check.
What should I keep after the agent says done?
Keep the file diff, exact command output, expected value, limitations and the original backup. Run the tests independently; prose alone is not verification.
Can the model’s instruction to be careful replace a sandbox?
No. A prompt expresses intent. The executor, process permissions and any configured OS boundary determine what the task can actually access.
Continue learning
Continue from the verified evidence
Compare published artifacts or return to the installation documentation.