Founder’s Corner · Evaluation Archive

Evaluation work,
in full.

This archive sits inside Founder’s Corner and presents completed evaluation work in one continuous collection: document-grounded reasoning, end-to-end software builds, and adversarial multi-source research.

One archive, complete work. Twenty finished evaluation tasks are presented together without client, community, or engagement names. The work spans document-grounded reasoning, end-to-end agentic software builds, and adversarial multi-source research.

What “in full” means. The document tasks retain prompts, critical elements, trajectories, rubrics, scoring, and ground truth. The software tasks include the complete brief, authored solution code, public verification code, Docker environment, evaluation contract, and design rationale. The research tasks retain the complete submission and testing record.

Code disclosure. The technical builds expose 72 authored Python, shell, and Docker files totaling 17,806 lines in an interactive source browser. Only project-owned names and project-specific identifiers are redacted; implementation logic and test behavior are not rewritten.

Privacy and provenance. Synthetic evaluation material is shown without private source documents, credentials, hidden verifier fixtures, customer data, or third-party project names.

Complete Evaluation & Build Archive · 20 tasks

Task 01

Credit Decision & Goal Planning

Tax planning · Credit decision · Goal-based planning

A 25-year-old analyst asks whether to attack a 9% Grad PLUS loan or a credit card, and whether his self-filed return is right.

6 documents

The PromptWhat the simulated consumer asks, in their own voice+

Persona 34 — Kwame Adjei: Output-Driven Human Prompt

Persona voice: 25-year-old single Chicago investment-banking analyst, high income, high debt, anxious but trying to get his life together. He writes like he talks — direct, a little scattered, occasionally self-deprecating.

Assigned categories: Credit decision, Goal-based planning


The Prompt

Hey — so I did my own taxes for 2025 and I think I owe the IRS like $4,700. Can you look at it and tell me if I fucked it up? I really don't want to pay a CPA if I don't have to.
Also I need help with something bigger. I make $110k base at Halstead Crane now, plus a $25k signing bonus, and there's a discretionary bonus too. I feel like I should be rich. Instead I have $160k in student loans, a credit card that won't die, and I'm trying to buy a condo in the next two years. On paper I'm doing great. In reality I'm a mess.
I want to know where my money actually goes. I know I spend too much but I don't know what to cut. Like, what is actually luxury and what just feels normal because I'm exhausted from work? Please go through my checking account and credit card and split everything into categories — rent, loans, food, subscriptions, stupid shit I don't need, all of it.
The other thing that scares me is this credit card. I set it on autopay so I wouldn't have to think about it, but every month the balance looks... higher? Not lower. That can't be right. I need to understand what's happening and what to do first — the card, the student loans, or saving for the condo. My gut says pay the student loans because $160k is huge, but the card feels more urgent somehow.
My hard rules: keep at least $6,000 cash between checking and savings, don't sell my VTI in the brokerage, don't mess with my 401(k), and don't tell me to hire a CPA. Just tell me what you'd do if you were me.
Here's what I can share:
- W2_Northbridge_2025.pdf — my old job Jan–May
- W2_HalsteadCrane_2025.pdf — new job Jun–Dec, has my 401(k) deferral
- 1099INT_Lakeshore_2025.pdf — interest income
- tax_tracker_DRAFT_2025.xlsx — my self-prepared tax spreadsheet
- offer_letter_HalsteadCrane.pdf
- paystub_HalsteadCrane_2026.pdf
- bank_checking_2026.pdf
- savings_statement_Lakeshore_2026.pdf
- brokerage_statement_Fidelity_2026.pdf
- student_loan_statement_2026.pdf
- credit_card_statement_2026.pdf
- transactions_checking_2026.csv — all my 2026 checking transactions
Walk me through it like I'm smart but tired. Start with the tax thing, then show me where the money is leaking, then tell me what to kill first and how long until I can realistically look at condos.

Required Output — give me these exact sections

Do not write a generic essay. I need tables, numbers, and a clear order of operations. Cite the file and line/field for every dollar figure.

1. Tax check (2025 federal only)

Open both W-2s, the 1099-INT, and my draft spreadsheet. Show me:

  • Combined Box 1 wages and combined Box 3/5 Social Security wages.
  • Why they are different (hint: 401(k)).
  • Whether I overpaid Social Security tax because of the $176,100 wage base. If yes, the exact credit I get back.
  • Whether I owe Additional Medicare Tax. If yes, the exact amount.
  • Whether I can really take the $2,500 student-loan interest deduction I put in my draft.
  • The corrected federal balance due (or refund) and how far off my $4,700 draft is.

I want the arithmetic visible, not just a final number. And I don't need legal advice — just tell me if my draft is wrong.

2. Where my money goes — spending autopsy table

Read transactions_checking_2026.csv and credit_card_statement_2026.pdf. Build a table with these exact categories and give me the 6-month total and monthly average for each:

Category What belongs here 6-month total Monthly avg
Fixed essentials Rent share, student loan payment, utilities, health
Variable essentials Groceries, transit
Subscriptions Netflix, Spotify, LA Fitness, DashPass, NYT, iCloud, etc. (list each separately)
Dining / delivery Restaurants, coffee, Uber Eats, DoorDash
Shopping / entertainment Amazon, Target, REI, Best Buy, Kindle, hobby stuff
Cash / fees ATM withdrawals, checking maintenance fee
Debt service Credit-card autopay — this is NOT spending
Savings / investing transfers Auto-transfer to savings, auto-deposit to brokerage — this is NOT spending

Call out anything that made you think "Kwame is paying for the same thing twice."

3. Luxury / stupid-shit / redundant list

Give me at least 5 specific cuts. For each one:

  • What it is.
  • Monthly cost.
  • Why it's redundant, luxury, or unnecessary.
  • Approximate monthly savings if I cut or reduce it.

Do not call rent, groceries, health, or my CTA pass "luxury." Be realistic — I work long hours, so some convenience spending is understandable, but some of it is just lazy.

4. The credit-card autopsy

I need to understand why my autopay isn't working. Use the statement to show me:

  • Opening balance, payments made, interest charged, purchases, closing balance over the 6-month period.
  • Monthly interest vs. minimum payment vs. new purchases.
  • The statement's projected payoff timeline if I only keep paying the minimum.
  • Why the balance is growing even though I'm "paying it."

Then tell me: pay this first, or pay the 9% Grad PLUS first? Show the math, not just a rule of thumb.

5. Monthly cash-flow math

Use my 2026 paystub to compute my monthly take-home pay. Then show me:

  • Fixed obligations after your recommended cuts.
  • Discretionary spending after your recommended cuts.
  • How much cash I can free up by pausing the savings auto-transfer and brokerage auto-deposit until the credit card is gone.
  • How many months it takes to pay off the credit card if I throw $500/mo, $1,000/mo, and $1,500/mo at it.

Make sure I never drop below $6,000 combined in checking + savings in any scenario.

6. Two-year condo-readiness plan

I want to shop for a condo in about 24 months. Show me the staged plan in order:

  1. Kill the credit card.
  2. Stabilize the $6,000 cash floor.
  3. Redirect the former card/savings/brokerage cash flow to the highest-interest student loan.
  4. Only then open a separate down-payment fund.

Also be honest: is a Chicago condo actually realistic in 24 months? Use my DTI, credit utilization, and down-payment math. If the answer is "only if you fix X," tell me exactly what X is.

7. Executive summary at the top

Before all the tables, give me a short paragraph answer to: "What is the one thing I should do this week?" Then a second paragraph for "What should I do this month?"


Hard Rules (do not break these)

  • Keep at least $6,000 combined in checking + savings at all times.
  • Do not sell my VTI brokerage holding.
  • Do not stop or reduce my 401(k) deferrals.
  • Do not frame this as licensed tax, legal, or investment advice.
  • Prioritize the 24.99% credit card over the 9% Grad PLUS loan.
  • Cite the specific file and field/line for every major number.
Critical ElementsThe findings a correct response must reach+

Persona 34 — Kwame Adjei: Critical Elements (Attempt 2)

Please describe the 10-15 most critical elements that the model should take into account when the prompt is solved or calculated.

These should be 2-3 sentences describing the critical elements in detail.

These elements should map directly to the correct answer and to your rubric. Each element should be verifiable and explicit.

1. The corrected 2025 federal balance due is $3,494.17, which is $1,208.83 lower than Kwame"s draft tracker claims.

Kwame"s draft tracker says he owes about $4,703, but that number is wrong once you reconcile both W-2s, remove the disallowed student-loan deduction, add the Additional Medicare Tax, and claim the excess Social Security credit from changing jobs mid-year. The model needs to walk through the actual 1040 chain and show the corrected bottom line, then state exactly how far the draft was off.

Rubric mapping: Corrected federal tax balance due and adjustment vs. draft

Canonical value / classification: Corrected balance due = $3,494.17; draft was $4,703.00; adjustment = −$1,208.83

Verdict trigger: PASS if the response shows $3,490–$3,500 and explains the $1,205–$1,215 downward adjustment from the draft; FAIL if it rubber-stamps the draft, lands outside the band, or cannot show the exact dollar difference.

2. The 2025 standard deduction is $15,750, the student-loan interest deduction is $0 because MAGI clears the phase-out, and the two-job situation produces a $1,864.13 excess Social Security credit plus $55.50 of Additional Medicare Tax.

The model has to use the 2025 single standard deduction of $15,750, recognize Kwame makes too much for any student-loan interest deduction because his MAGI is about $204,787, and handle the two-job quirks: Social Security tax withheld on $206,167 in Box 3 wages exceeds the $176,100 wage base, so the over-withholding becomes a refundable credit, while the Medicare wages over $200,000 trigger a small Additional Medicare Tax owed. Each of these moves the final number, and skipping any one breaks the reconciliation.

Rubric mapping: Tax mechanics: standard deduction, SL deduction phase-out, excess SS credit, and Additional Medicare Tax

Canonical value / classification: Standard deduction = $15,750; allowed SL deduction = $0; excess SS credit = $1,864.13; Additional Medicare Tax = $55.50

Verdict trigger: PASS if all four items are stated with their correct values and the disallowed deduction is tied to MAGI over $100,000; FAIL if the SL deduction is kept at $2,500, the excess SS credit is missed, or the Additional Medicare Tax is omitted.

3. Monthly net take-home pay from the 2026 paystub is $5,852.24, not a rough estimate off the $110,000 salary.

Kwame thinks his take-home should be bigger because he earns $110,000 base, but the paystub is what matters: after medical and 401(k) pre-tax deductions and all withholdings, each semimonthly check nets $2,926.12. The model must derive the monthly take-home from that stub, not from a rough estimate off the $110,000 salary.

Rubric mapping: Monthly net take-home pay calculation

Canonical value / classification: $5,852.24 per month

Verdict trigger: PASS if the response derives monthly net pay around $5,850–$5,855 from the semimonthly paystub; FAIL if it guesses from gross salary, ignores pre-tax deductions, or misses the paystub as the source.

4. True monthly consumption is roughly $2,260–$2,350, after separating debt service and internal transfers from spending.

The model cannot just add up every outflow from the checking statement and call it spending. The $1,847 student-loan payment and the roughly $63 credit-card autopay are debt service, while the $300 savings auto-transfer and $250 brokerage auto-deposit are internal transfers to Kwame"s own accounts; only the remaining categories like rent, groceries, subscriptions, dining, shopping, ATM cash, and fees are actual lifestyle consumption.

Rubric mapping: Spending autopsy and transfer/debt-service classification

Canonical value / classification: True monthly consumption ≈ $2,260–$2,350; debt service = $1,847 + ~$63; internal transfers = $300 + $250

Verdict trigger: PASS if the response separates those four non-consumption flows from spending and lands in the consumption band; FAIL if any of those four flows are counted as spent or if total consumption is inflated by double-counting.

5. The credit card is negatively amortizing, and the statement"s minimum-payoff warning is 23 years, 4 months / $12,744.61.

Kwame set the card to autopay minimums, yet the balance keeps climbing because the average monthly interest of about $65–$70 is larger than the average autopay of about $60–$67, so even without new purchases the principal would drift upward. The model must explain that mechanic clearly and quote the statement"s minimum-payoff projection exactly, not round it off.

Rubric mapping: Credit-card mechanics and negative-amortization diagnosis

Canonical value / classification: Balance = $3,383.83; APR = 24.99%; avg monthly interest ≈ $65.59; avg monthly payment ≈ $63.08; minimum-payoff timeline = 23 years, 4 months; minimum-payoff total cost = $12,744.61

Verdict trigger: PASS if the response explains interest exceeds payments and quotes both 23 years, 4 months and $12,744.61; FAIL if it only says "pay more" without the negative-amortization explanation or misses either quoted value.

6. Paying an extra $500 per month clears the card in about 7 months; an extra $1,000 clears it in about 4 months.

Once Kwame sees the 23-year trap, he needs hard numbers on how fast extra cash kills the card. The model should amortize the $3,383.83 balance at 24.99% APR using the minimum autopay plus the extra amount and show the payoff timelines for both the $500 and $1,000 monthly boosts he asked about.

Rubric mapping: Credit-card payoff acceleration scenarios

Canonical value / classification: +$500/month → ~7 months; +$1,000/month → ~4 months

Verdict trigger: PASS if both timelines fall within ±1 month of the canonical values; FAIL if only one scenario is computed, the amortization ignores the existing minimum, or either timeline is off by more than one month.

7. The highest-cost debt is the 24.99% credit card, followed by Group C Grad PLUS at 9.00%, then Group D at 7.12%, Group B at 6.54%, and Group A at 4.99%.

Kwame"s instinct is to attack the $160,000 total, but the model needs to break the loans into the four groups and rank them by rate, not balance. Group C Grad PLUS is the worst student-loan bucket at 9.00%, but the credit card at 24.99% is the highest-cost debt overall, so every extra dollar should go to the card first and only later to the Grad PLUS.

Rubric mapping: Student-loan stack identification and debt-priority reasoning

Canonical value / classification: Group A $40,610.25 @ 4.99%; Group B $56,080.83 @ 6.54%; Group C $45,444.80 @ 9.00%; Group D $17,404.38 @ 7.12%; highest-rate student group = Group C; priority = credit card first, then Grad PLUS Group C

Verdict trigger: PASS if all four groups are listed with correct balances and rates, Group C is identified as the worst student loan, and the response says the 24.99% card must be paid before any student-loan prepayment; FAIL if the highest-rate group is misidentified or if extra payments are directed to loans while the card still carries a balance.

8. Non-housing back-end DTI is about 20.9% and credit utilization is about 42.3%; utilization must drop under 10% for a mortgage.

A mortgage lender will look at Kwame"s $9,166.67 monthly gross from the $110,000 base salary and compare it to his recurring non-housing debt of $1,847 student loans plus about $67 minimum card payment, giving a non-housing back-end DTI around 20.9%. The model also needs to flag the 42.3% credit utilization on the $8,000 limit card and explain that getting under 10% is a prerequisite, while lowering the Grad PLUS balance would free DTI headroom once PITIA is added.

Rubric mapping: Mortgage-readiness DTI and credit-utilization analysis

Canonical value / classification: Non-housing DTI = 20.89%; credit utilization = 42.30%; utilization target = under 10%

Verdict trigger: PASS if DTI is in the 20.8%–21.0% band, utilization is in the 42.0%–43.0% band, the target utilization is under 10%, and the response explains both must improve; FAIL if it uses net pay or bonus income for DTI, omits utilization, or ignores what a lender would require.

9. Redirecting $550 per month to 1.25% savings/brokerage while owing 24.99% on the card loses roughly 23.7 percentage points of annualized spread.

Kwame is automatically sending $300 to a savings account earning about 1.25% and $250 to a brokerage while he owes 24.99% on the credit card, which is a guaranteed losing spread of roughly 23.7 percentage points on that $550 every month. The model should quantify that spread and recommend pausing those auto-flows until the card is zero.

Rubric mapping: Negative arbitrage / self-defeating cash-flow pattern

Canonical value / classification: $550/month redirected at ~1.25% while owing 24.99%; annualized loss spread ≈ 23.7 percentage points

Verdict trigger: PASS if the response quantifies the ~23–24 percentage point spread and says the savings/brokerage auto-flows should pause until the card is paid off; FAIL if it leaves those flows unchanged, calls them good habits without the math, or fails to quantify the spread.

10. Specific keep / cancel / reduce actions should free roughly $210–$297 per month without touching essentials, the 401(k), or VTI.

The model needs to give Kwame concrete, defensible actions he can take right now, not vague "spend less" advice. That means canceling DoorDash DashPass because he pays $9.99/mo for it but shows no DoorDash food orders, while Uber Eats appears on the credit card every month; switching or waiving the checking maintenance fee, cutting ATM cash withdrawals, trimming dining and delivery, and reducing discretionary shopping, while explicitly protecting rent, groceries, health, transit, the 401(k), and the VTI holding.

Rubric mapping: Keep / cancel / reduce spending recommendations

Canonical value / classification: Cancel DashPass ≈ $9.99/mo; avoid checking fee ≈ $10–$12/mo; reduce ATM cash ≈ $40–$50/mo; reduce dining/delivery ≈ $50–$75/mo; trim shopping ≈ $100–$150/mo; total realistic freed cash ≈ $210–$297/mo

Verdict trigger: PASS if at least five quantified actions are given, DashPass is flagged as an unused subscription while Uber Eats is actively used, and essentials plus 401(k) and VTI are protected; FAIL if the advice is generic, mislabels an essential as cuttable, or recommends a CPA, VTI sale, or 401(k) reduction.

11. The 2-year plan has three phases — immediate tax/card payoff, a 3-month buffer rebuild, then 21 months of Grad PLUS/down-payment splitting — and preserves the $6,000 cash floor.

The plan has to start with paying the corrected $3,494 tax and the $3,383.83 card balance out of the roughly $20,862 in combined cash, leaving well over the $6,000 floor. Then it should pause the $300 savings and $250 brokerage auto-flows for a few months while applying the spending cuts to rebuild the cash buffer, and only after that split freed cash between attacking the 9.00% Grad PLUS and building a down-payment fund, all while respecting the no-CPA, no-VTI-sale, no-401(k)-reduction, and $6,000-floor rules.

Rubric mapping: Sequenced 2-year condo-readiness plan

Canonical value / classification: Immediate: tax + card paid from liquid cash, remaining cash ≈ $13,984 (≥ $6,000); Months 1–3: pause $550 auto-flows + apply cuts to rebuild buffer; Months 4–24: redirect freed cash to Grad PLUS prepayment and down-payment fund

Verdict trigger: PASS if the plan has phase-by-phase month counts, shows the cash floor is preserved, keeps VTI and 401(k) untouched, does not recommend a CPA, and moves to Grad PLUS only after the card is zero; FAIL if any hard constraint is violated, the timeline is missing, or the plan jumps to student-loan prepayment before the card is gone.

12. Every major number should be tied to a named source document in the workspace.

Kwame explicitly asked which document each number came from, so the model should trace the tax figures back to the two W-2s, the 1099-INT, and the draft tracker; the paystub for take-home; the checking and credit-card statements for spending; the credit-card statement for balance, APR, minimum payment, and payoff warning; the student-loan statement for the four groups; and the bank, savings, and brokerage statements for cash and the auto-flows. Citations do not need to be footnotes, but the source document should be named when each key figure is introduced.

Rubric mapping: Document sourcing and traceability

Canonical value / classification: Each major figure tied to a named source document in the workspace

Verdict trigger: PASS if the response names the source document for the tax chain, paystub, spending categories, card mechanics, loan groups, DTI/utilization inputs, and cash positions; FAIL if numbers appear without attribution or the model says the documents are unavailable.

13. The condo verdict is conditional: realistic only if the card is zero, utilization is under 10%, Grad PLUS is materially reduced, and a down-payment fund is built.

The model has to give Kwame an honest answer: buying a Chicago condo in two years is not impossible, but it is conditional. The verdict should say it only becomes realistic if the card is zero, utilization drops under 10%, the Grad PLUS balance is materially reduced to free DTI headroom, and a real down-payment fund is built, given that the current 20.9% non-housing DTI leaves limited room once PITIA is added.

Rubric mapping: Condo goal feasibility verdict

Canonical value / classification: Conditional / realistic only if card is zero, utilization <10%, Grad PLUS reduced, and down-payment fund built

Verdict trigger: PASS if the response gives a clear conditional verdict tied to those four conditions and does not declare the goal fully on-track or impossible without explanation; FAIL if it says the condo is definitely achievable now, or rules it out without explaining what would make it achievable.

Golden TrajectoryStep-by-step path to the answer, every figure sourced+

Persona 34 — Kwame Adjei: Golden Trajectory (Attempt 2)

Step-by-step path from the prompt to the verified answer

1. Reconstruct 2025 wages and interest income
  • Navigate to W2_Northbridge_2025.pdf and read Box 1 wages $38,000.00 and Box 3 Social Security wages $38,000.00.
  • Navigate to W2_HalsteadCrane_2025.pdf and read Box 1 wages $164,166.67, Box 3 Social Security wages $168,166.67, and Box 4 Social Security tax withheld $10,426.33.
  • Navigate to 1099INT_Lakeshore_2025.pdf and read Box 1 taxable interest $120.00.
  • Calculate combined Box 1 wages: $38,000.00 + $164,166.67 = $202,166.67.
  • Calculate combined Box 3 Social Security wages: $38,000.00 + $168,166.67 = $206,166.67.
  • Calculate AGI: $202,166.67 + $120.00 = $202,286.67.
2. Verify the draft"s student-loan interest deduction is disallowed
  • Navigate to tax_tracker_DRAFT_2025.xlsx and confirm the draft claims a $2,500.00 student-loan interest deduction and a draft balance due of $4,703.00.
  • Calculate MAGI for the student-loan interest deduction: $202,286.67 + $2,500.00 = $204,786.67.
  • Apply the 2025 single-filer phase-out ceiling of $100,000.00; because MAGI exceeds the ceiling, the allowed deduction is $0.00.
3. Compute 2025 taxable income and regular tax
  • Use the 2025 single standard deduction $15,750.00.
  • Calculate taxable income: $202,286.67 − $15,750.00 − $0.00 = $186,536.67.
  • Apply 2025 single brackets:
  • 10% on $11,925.00 = $1,192.50
  • 12% on $36,550.00 = $4,386.00
  • 22% on $54,875.00 = $12,072.50
  • 24% on $83,186.67 = $19,964.80
  • Sum to regular income tax: $37,615.80.
4. Compute excess Social Security tax credit and Additional Medicare Tax
  • Sum total Social Security tax withheld from both W-2s Box 4: $2,356.00 (Northbridge) + $10,426.33 (Halstead) = $12,782.33.
  • Use the 2025 Social Security wage base $176,100.00 and calculate maximum Social Security tax: $176,100.00 × 6.2% = $10,918.20.
  • Calculate excess Social Security credit: $12,782.33 − $10,918.20 = $1,864.13.
  • Use combined Medicare wages (Box 5) $206,166.67 and the 2025 Additional Medicare Tax single threshold $200,000.00.
  • Calculate Additional Medicare Tax: ($206,166.67 − $200,000.00) × 0.9% = $55.50.
5. Compute corrected 2025 federal balance due and adjustment vs. draft
  • Sum total corrected tax liability: $37,615.80 + $55.50 − $1,864.13 = $35,807.17.
  • Sum combined federal withholding from both W-2s Box 2: $4,200.00 + $28,113.00 = $32,313.00.
  • Calculate corrected balance due: $35,807.17 − $32,313.00 = $3,494.17.
  • Calculate adjustment from draft: $3,494.17 − $4,703.00 = −$1,208.83.
6. Compute monthly net take-home pay from the 2026 paystub
  • Navigate to paystub_HalsteadCrane_2026.pdf and read semimonthly gross pay $4,583.33, medical pre-tax deduction $285.00, 401(k) deferral $230.00, and net pay $2,926.12.
  • Calculate monthly net take-home: $2,926.12 × 2 = $5,852.24.
7. Separate actual spending from debt service and internal transfers
  • Navigate to transactions_checking_2026.csv and identify the four non-consumption flows:
  • GREAT LAKES STUDENT LOANS PMT = $1,847.00/month (debt service)
  • MERIDIAN BANK CARD AUTOPAY = average $63.08/month over 6 months (debt service)
  • TRANSFER TO SAVINGS *LAKESHORE = $300.00/month (internal transfer)
  • TRANSFER TO BROKERAGE *FIDELITY = $250.00/month (internal transfer)
  • Sum the remaining categories (rent, utilities, groceries, dining, transit, subscriptions, shopping, ATM cash, fees) over 6 months and divide by 6 to get true monthly consumption of approximately $2,333/month, within the $2,260–$2,350 band.
  • Flag DOORDASH *DASHPASS at $9.99/month as an unused subscription; verify no DoorDash food orders appear while UBER EATS CHICAGO IL appears monthly on the credit card.
8. Diagnose the credit-card negative amortization
  • Navigate to credit_card_statement_2026.pdf and read new balance $3,383.83, credit limit $8,000.00, APR 24.99%, minimum payment due $67.68, and the Minimum Payment Warning: 23 years, 4 months and total cost $12,744.61.
  • Sum total interest charged year-to-date $393.54 and total payments $378.48 from the Year-to-Date Totals.
  • Calculate average monthly interest: $393.54 ÷ 6 = $65.59/month.
  • Calculate average monthly payment: $378.48 ÷ 6 = $63.08/month.
  • Show that because average interest ($65.59) exceeds average payment ($63.08), the principal drifts upward even without new purchases.
  • Calculate payoff acceleration:
  • Extra $500/month (total ~$567.68/month) → about 7 months.
  • Extra $1,000/month (total ~$1,067.68/month) → about 4 months.
9. Build the student-loan stack and debt priority
  • Navigate to student_loan_statement_2026.pdf and read the four groups:
  • Group A: $40,610.25 @ 4.99%
  • Group B: $56,080.83 @ 6.54%
  • Group C Grad PLUS: $45,444.80 @ 9.00%
  • Group D: $17,404.38 @ 7.12%
  • Identify Group C as the highest-rate student-loan group at 9.00%.
  • Rank all debt by rate: credit card 24.99% first, then Group C 9.00%, Group D 7.12%, Group B 6.54%, Group A 4.99%.
10. Compute lender-style DTI and credit utilization
  • Navigate to offer_letter_HalsteadCrane.pdf or paystub_HalsteadCrane_2026.pdf for base salary $110,000/year.
  • Calculate gross monthly income: $110,000 ÷ 12 = $9,166.67/month.
  • Sum non-housing debt payments: $1,847.00 (student loans) + $67.68 (credit card minimum) = $1,914.68/month.
  • Calculate non-housing back-end DTI: $1,914.68 ÷ $9,166.67 = 20.89%.
  • Calculate credit utilization: $3,383.83 ÷ $8,000.00 = 42.30%.
  • State lender target: utilization under 10%.
11. Quantify negative arbitrage from savings/brokerage flows
  • Navigate to savings_statement_Lakeshore_2026.pdf and confirm the savings APY is about 1.25% and the auto-transfer is $300.00/month.
  • Navigate to brokerage_statement_Fidelity_2026.pdf and confirm the auto-deposit is $250.00/month.
  • Calculate total auto-flow while credit card revolves: $300.00 + $250.00 = $550.00/month.
  • Calculate annualized loss spread: 24.99% (card APR) − 1.25% (savings yield) ≈ 23.7 percentage points.
  • Recommend pausing both auto-flows until the credit-card balance is zero.
12. List specific keep / cancel / reduce actions
  • From transactions_checking_2026.csv, identify:
  • Cancel DOORDASH *DASHPASS = $9.99/month (unused subscription).
  • Waive/eliminate MONTHLY MAINTENANCE FEE = $10.00–$12.00/month.
  • Reduce ATM cash withdrawals from ~$83/month to ~$35/month = $40.00–$50.00/month.
  • Reduce dining/delivery from ~$154/month to ~$90/month = $50.00–$75.00/month.
  • Reduce discretionary shopping from ~$290/month to ~$170/month = $100.00–$150.00/month.
  • Sum realistic freed cash: approximately $210.00–$297.00/month.
  • Explicitly protect rent, groceries, health, transit, 401(k), and VTI.
13. Build the sequenced 2-year condo-readiness plan
  • Navigate to bank_checking_2026.pdf and savings_statement_Lakeshore_2026.pdf and confirm combined liquid cash of approximately $20,862.00.
  • Immediate (Month 1): Pay corrected tax $3,494.17 and credit card $3,383.83 from cash.
  • Calculate remaining cash: $20,862.00 − $3,494.17 − $3,383.83 = $13,984.00, which exceeds the $6,000.00 floor.
  • Months 1–3: Pause $550.00/month in savings/brokerage auto-flows and apply ~$210.00–$297.00/month in spending cuts to rebuild the cash buffer.
  • Months 4–24: Split freed cash between extra principal payments on Group C Grad PLUS (9.00%) and a dedicated down-payment fund, targeting a down payment of $30,000.00–$40,000.00 by Month 24.
  • Months 20–24: Maintain zero credit-card balance (utilization ~0%), reduce Grad PLUS balance materially to free DTI headroom, and establish two-year bonus history so lenders can consider bonus income.
14. Deliver the conditional condo feasibility verdict
  • State the verdict: a Chicago condo in roughly two years is conditional, not guaranteed.
  • List the four conditions:
  1. Credit-card balance is zero and utilization is under 10%.
  2. Grad PLUS balance is materially reduced to free DTI room.
  3. Down-payment fund of ~$30,000–$40,000 is built.
  4. $6,000 cash floor is preserved, 401(k) is untouched, and VTI is not sold.
  • Cite the source for the constraint: current 20.89% non-housing DTI leaves limited room once PITIA is added.

Summary of verified outputs

Output Verified value
Corrected 2025 federal balance due $3,494.17
Adjustment vs. draft $4,703 −$1,208.83
Excess Social Security credit $1,864.13
Additional Medicare Tax $55.50
Monthly net take-home pay $5,852.24
True monthly consumption ~$2,260–$2,350
Credit-card balance / APR / utilization $3,383.83 / 24.99% / 42.30%
Statement minimum-payoff warning 23 years, 4 months / $12,744.61
+$500 payoff / +$1,000 payoff ~7 months / ~4 months
Non-housing back-end DTI 20.89%
Realistic monthly cuts ~$210–$297
Remaining cash after tax + card payoff ~$13,984
Scoring RubricWeighted positive and negative criteria+

Persona 34 (Kwame Adjei) — Attempt 2 Grading Rubric

Prompt: persona_34_kwame_adjei_attempt2_prompt.md

Critical elements: persona_34_attempt2_critical_elements.md

Golden trajectory: persona_34_attempt2_golden_trajectory.md

Rationale/Failures: persona_34_attempt2_failure_justifications.md

Rubric summary

  • 35 criteria total, all numeric IDs
  • Positive weights sum to +730
  • Negative weights sum to -320
  • Headline criteria: #1, #2
  • Deal-breaker negative: #31, #33

Scoring model

Tier Range Purpose
Headline +80 to +100 Single most important correct outcome
Important +25 to +55 Key explicit asks / drivers of the plan
Minor +10 to +20 Fine-grained verification / supporting details
Deal-breaker -60 to -100 Fatal commission errors
Major -30 to -55 Serious commission errors
Moderate -10 to -25 Tangible but less harmful errors

Positive criteria

ID Criterion Weight Rationale
1 Response states the corrected 2025 federal balance due is $3,494.17 within +/- $10 80 This is the headline number the user asked to verify; it is only correct after the two-job SS and Medicare reconciliation.
2 Response states the adjustment from the draft $4,703 is a downward $1,208.83 within +/- $15 30 The prompt explicitly asks for the exact dollar adjustment from the draft; getting the direction right is as important as the final balance.
3 Response uses the 2025 single standard deduction of $15,750 20 Using $15,000 inverts the correct 2025 single-filer constant and corrupts the tax chain.
4 Response states the student-loan interest deduction is $0 because MAGI exceeds the $100,000 phase-out ceiling 20 The draft wrongly kept the $2,500 deduction; MAGI of about $204,800 clears the ceiling.
5 Response computes the $1,864.13 excess Social Security withholding credit from the two W-2s Box 4 total above the $176,100 wage base 40 This two-employer credit is the largest downward driver of the corrected balance.
6 Response computes the $55.50 Additional Medicare Tax on combined Medicare wages above $200,000 25 All three responses omitted this tax, so their corrected balances were wrong.
7 Response derives monthly net take-home as $5,852.24 from the semimonthly paystub 25 The prompt asks for paystub-based cash-flow math, not a rough estimate off the $110k base.
8 Response classifies the $300 savings auto-transfer, $250 brokerage auto-deposit, $1,847 student-loan payment, and ~$63 credit-card autopay as non-consumption flows rather than monthly spending 40 The prompt explicitly asks to separate these four flows from actual consumption; Responses 1 and 2 mislabeled them as spending or checking outflow.
9 Response reports true monthly consumption in the $2,260-$2,350 band and sorts spending into categories such as rent, groceries, dining, transit, subscriptions, shopping, ATM cash, and fees 25 The band confirms no double-counting and the categories match the prompt's requested spending autopsy.
10 Response states the credit-card new purchases total for the period is $370.25 15 The prompt explicitly asks for interest versus minimum payment versus new purchases.
11 Response explains the credit-card balance grows because monthly interest exceeds the minimum autopay, and quotes the minimum-payoff warning as 23 years, 4 months and $12,744.61 35 The negative-amortization mechanism and the exact statement warning are both required.
12 Response states the average monthly interest charge on the credit card is $65.59 within +/- $5 15 Computed from $393.54 year-to-date interest over 6 months; this is one third of the prompt's interest-vs-minimum-vs-purchases ask.
13 Response states the credit-card minimum payment is $67.68 10 This is the second third of the prompt's interest-vs-minimum-vs-purchases ask.
14 Response states that paying an extra $500/month clears the $3,383.83 balance in 7 months +/- 1 month 15 Response 3 said 8 months, which is the observed off-by-one failure.
15 Response states that paying an extra $1,000/month clears the $3,383.83 balance in 4 months +/- 1 month 15 The prompt explicitly asks for this acceleration scenario.
16 Response lists the four student-loan groups as Group A $40,610.25 @ 4.99%, Group B $56,080.83 @ 6.54%, Group C $45,444.80 @ 9.00%, and Group D $17,404.38 @ 7.12% 25 The four-group breakdown is required for correct debt-priority and payoff sequencing.
17 Response identifies the 24.99% credit card as the highest-cost debt 20 This is one half of the prompt's credit-card-first versus student-loans question.
18 Response identifies Grad PLUS Group C at 9.00% as the highest-cost student loan 20 This is the second half of the prompt's debt-priority question.
19 Response computes non-housing back-end DTI as $1,914.68 divided by $9,166.67, giving 20.9% +/- 0.5% 20 The prompt asks for lender-style DTI, which uses base gross and excludes the bonus.
20 Response states credit utilization is $3,383.83 / $8,000.00 = 42.3% +/- 0.5% and the mortgage-readiness target is under 10% 20 Both the current ratio and the target are required for the condo-readiness verdict.
21 Response quantifies the negative arbitrage as 23.7 percentage points +/- 1 percentage point on $550/month sent to savings/brokerage while the 24.99% credit card revolves 30 All three responses kept the auto-flows running despite the 24.99% card versus ~1.25% savings spread.
22 Response recommends pausing the $300 savings auto-transfer and $250 brokerage auto-deposit until the credit-card balance is zero 25 This is the concrete action required to stop the self-defeating spread.
23 Response identifies at least five specific keep/cancel/reduce cuts with quantified monthly savings 20 Response 3's cut list was too soft; the prompt explicitly asks for a keep/cancel/reduce list with dollar savings.
24 Response cancels DoorDash DashPass because it shows six $9.99 subscription charges and no DoorDash food orders while Uber Eats appears monthly 15 DashPass is paid but unused while Uber Eats is actively charged.
25 Response flags the monthly checking maintenance fee of $10-$12 as avoidable 10 This is directly visible in the checking CSV and is a high-confidence cut.
26 Response presents a sequenced 2-year plan with phase month counts 40 The prompt explicitly asks for how many months each step takes.
27 Response sequences the plan as corrected tax and credit card first, then Grad PLUS, then a down-payment fund 30 This ordering is the core strategic insight of the plan.
28 Response names the source document for each figure in the tax chain, paystub, spending, credit card, student loans, DTI/utilization, and cash positions 20 The prompt explicitly asks to name the document each number came from.
29 Response gives a conditional condo verdict tied to: credit card zero and utilization under 10%, Grad PLUS balance reduced by at least $8,000 by Month 24, down-payment fund of $30,000-$40,000 by Month 24, and $6,000 cash floor preserved with VTI and 401(k) untouched 25 The prompt asks for an honest feasibility verdict on the two-year Chicago condo goal.

Negative criteria

ID Criterion Weight Rationale
30 Response recommends extra payments to any student-loan group while the 24.99% credit card still carries a balance -50 Any extra payment to a lower-rate loan while the card carries a balance is mathematically wrong and contradicts the headline priority.
31 Response presents the draft $4,703 federal balance as the final corrected answer -80 Restating the draft number leaves the user with the original wrong tax surprise.
32 Response labels the $300 savings transfer, $250 brokerage deposit, $1,847 student-loan payment, or ~$63 credit-card autopay as ordinary spending or checking outflow -45 Responses 1 and 2 committed this exact classification error, inflating the spending bottom line.
33 Response recommends selling VTI, reducing or stopping 401(k), dropping below $6,000 cash, or hiring a CPA -70 The prompt explicitly forbids all five of these escape hatches.
34 Response proposes cutting rent, groceries, health insurance, or CTA/transit by $25/month or more, or labels them luxury/major cut targets -35 The prompt treats these as essential needs; cuts above a token $25 undermine the plan's credibility and the user's stated values.
35 Response states a corrected federal balance outside $3,494.17 +/- $150 and presents it as a definitive figure -40 Response 3 produced a wrong tax headline and presented it as final, distorting downstream recommendations.

Model scores (projected)

Response Pass positives Hit negatives Projected raw Outcome
Response 1 ~80 pts (tax headline + a few partials) #30 (-50), #31 (-80), #32 (-45), #33 (-70) -165 Fail
Response 2 ~100 pts (tax headline + some pieces) #30 (-50), #31 (-80), #32 (-45), #33 (-70), #35 (-40) -185 Fail
Response 3 ~120 pts (better consumption work but still wrong tax) #30 (-50), #31 (-80), #33 (-70), #35 (-40) -120 Fail

All three responses fail because each either misses the headline corrected tax balance, recommends paying student loans while the 24.99% card revolves, violates the forbidden escape hatches, or misclassifies non-consumption flows as spending.

Verification notes

  • Corrected federal balance: $4,703.00 - $1,864.13 + $55.50 = $3,494.17
  • Excess SS credit: $11,187.00 (W-2 #1) + $1,595.33 (W-2 #2) = $12,782.33; $12,782.33 - $10,918.20 = $1,864.13
  • Additional Medicare Tax: ($110,500.00 + $94,286.67) - $200,000 = $4,786.67; × 0.009 = $55.50
  • Net take-home: semimonthly $2,926.12 × 2 = $5,852.24
  • True consumption: $5,852.24 - $300 savings - $250 brokerage - $1,847.00 SL - $63.08 CC autopay - $2,000 401(k) estimate = $2,392.16; further adjusted by bonus timing and deposits gives $2,260-$2,350
  • Credit card: $3,383.83 balance, $370.25 new purchases, $393.54 YTD interest over 6 months = ~$65.59/month, $67.68 minimum, $8,000 limit, warning 23 years 4 months / $12,744.61
  • Non-housing DTI: $1,914.68 / $9,166.67 = 20.89%
  • Utilization: $3,383.83 / $8,000.00 = 42.30%
Model ScoringHow each model response scored, and why+

Persona 34 — Kwame Adjei: Failure Justifications (Attempt 2)

Response 1

FAIL. Response 1 states: "used an incorrect standard deduction of $15,750 instead of the actual 2025 single deduction of $15,000," which inverts the verified 2025 single standard deduction of $15,750. It also omits the Additional Medicare Tax entirely, so its headline "corrected balance due" lands at $3,618.67 instead of the verified $3,494.17, and the adjustment versus the draft is off by $124.50. A reviewer cannot ship a response that reverses a basic 2025 tax constant and then mislabels the $1,847 student-loan payment as "Actual Monthly Spending," inflating true consumption from ~$2,333 to ~$4,120. The 2-year plan is also incomplete: it never attacks the 9.00% Grad PLUS group and instead routes freed cash only to "high-yield savings" while keeping the $300 savings and $250 brokerage auto-flows running, bypassing the ~23.7 percentage-point negative arbitrage.

Response 2

FAIL. Response 2 repeats the same tax-constant error, writing "Standard deduction … $15,000.00 — Draft used 2026 single amount, not 2025," and omits the $55.50 Additional Medicare Tax, producing a corrected balance of $3,619 rather than $3,494.17. It correctly labels the $1,847 student-loan payment as "Actual debt payment" in one table, but then folds it into "Your actual checking outflow excluding card autopay and transfers" of $4,156.31/month, conflating debt service with lifestyle spending and missing the verified true-consumption band of ~$2,260–$2,350. A reviewer would refuse to ship this because the headline spending bottom line is contradicted by the response"s own table, the tax chain is corrupted by the wrong standard deduction, and the plan never quantifies the negative-arbitrage spread or pauses the savings/brokerage auto-flows while the 24.99% card still carries a balance.

Response 3

FAIL. Response 3 tells Kwame: "Balance due ≈ $5,302" and "The exact adjustment from your draft: +$599," which is the opposite direction and magnitude of the verified answer — the actual corrected balance is $3,494.17 and the adjustment is −$1,208.83. It mentions the possibility of excess Social Security withholding but never applies the verified $1,864.13 credit, and it omits the $55.50 Additional Medicare Tax. The credit-card math is also off, stating "$500/mo … ~8 months" instead of the verified ~7 months, and the spending-cut list is too soft at "~$55–$95/mo" because it misses the larger ATM and shopping reductions visible in the checking CSV. A reviewer would reject this because the headline tax answer is materially wrong in both direction and dollars, and the 2-year plan never includes payment of the corrected tax bill or pauses the $300 savings and $250 brokerage transfers that are losing ~23.7 percentage points to the revolving 24.99% card.

Ground TruthVerified reference calculations+

Persona 34 — Kwame Adjei: Verified Ground-Truth Reference Sheet

Verifier scope and source-file inventory

This reference sheet is built strictly from the project documents in the workspace. No external IRS research, web searches, or outside knowledge beyond standard 2025 tax constants were used.

The persona prompt names 11 source files (PDFs, CSV, XLSX). None of those files are present in the workspace as standalone documents. The workspace contains only these Markdown files:

  • persona_34_kwame_adjei_human_prompt.md
  • persona_34_kwame_adjei_critical_elements.md
  • persona_34_kwame_adjei_golden_trajectory_formatted.md
  • persona_34_kwame_adjei_rubric_FINAL.md
  • persona_34_ground_truth_verification.md (this file)
  • persona_34_model_fail_justifications.md

All dollar figures below were extracted from the embedded reference tables in those Markdowns, then recomputed independently. For scoring consistency, this sheet cites the canonical source-file names named in the prompt.


1. 2025 federal tax chain — corrected balance due

1.1 Wage reconstruction
Item Verified value Source / formula
W2_Northbridge_2025.pdf, Box 1 $38,000.00 Embedded reference
W2_HalsteadCrane_2025.pdf, Box 1 $164,166.67 Embedded reference
Combined Box 1 wages $202,166.67 $38,000.00 + $164,166.67
W2_Northbridge_2025.pdf, Box 3 $38,000.00 Embedded reference
W2_HalsteadCrane_2025.pdf, Box 3 $168,166.67 Embedded reference
Combined Box 3 / Social Security wages $206,166.67 $38,000.00 + $168,166.67
401(k) deferral gap at Halstead $4,000.00 Box 3 − Box 1 = $168,166.67 − $164,166.67
1099INT_Lakeshore_2025.pdf, Box 1 $120.00 Embedded reference
AGI $202,286.67 Combined Box 1 wages + interest = $202,166.67 + $120.00
1.2 Standard deduction and taxable income
Item Verified value Source / formula
2025 single standard deduction $15,750.00 2025 IRS constant
Student-loan interest deduction (draft) $2,500.00 tax_tracker_DRAFT_2025.xlsx (draft claim)
MAGI for student-loan deduction $204,786.67 AGI + $2,500 = $202,286.67 + $2,500
2025 single phase-out ceiling $100,000.00 2025 IRS constant
Allowed student-loan deduction $0.00 MAGI $204,786.67 > $100,000
Taxable income $186,536.67 AGI − standard deduction − SL deduction = $202,286.67 − $15,750.00 − $0
1.3 Regular income tax (2025 single brackets)
Bracket Amount in bracket Rate Tax
$0 – $11,925 $11,925.00 10% $1,192.50
$11,926 – $48,475 $36,550.00 12% $4,386.00
$48,476 – $103,350 $54,875.00 22% $12,072.50
$103,351 – $197,300 $83,186.67 24% $19,964.80
Total regular tax $37,615.80

Formula check: $1,192.50 + $4,386.00 + $12,072.50 + $19,964.80 = $37,615.80.

1.4 Excess Social Security tax credit
Item Verified value Source / formula
W2_Northbridge_2025.pdf, Box 4 $2,356.00 Embedded reference
W2_HalsteadCrane_2025.pdf, Box 4 $10,426.33 Embedded reference
Total Social Security tax withheld $12,782.33 $2,356.00 + $10,426.33
2025 Social Security wage base $176,100.00 2025 IRS constant
Maximum Social Security tax $10,918.20 $176,100.00 × 6.2%
Excess Social Security credit $1,864.13 $12,782.33 − $10,918.20
Overage formula ($206,166.67 − $176,100.00) × 6.2% = $30,066.67 × 6.2% = $1,864.13
1.5 Additional Medicare Tax
Item Verified value Source / formula
Combined Medicare wages (Box 5) $206,166.67 Same as combined Box 3
2025 Additional Medicare Tax threshold (single) $200,000.00 2025 IRS constant
Overage $6,166.67 $206,166.67 − $200,000.00
Additional Medicare Tax $55.50 $6,166.67 × 0.9%
1.6 Corrected federal balance due
Item Verified value Source / formula
Regular income tax $37,615.80 Computed above
Additional Medicare Tax $55.50 Computed above
Excess Social Security credit −$1,864.13 Computed above
Total corrected federal tax liability $35,807.17 $37,615.80 + $55.50 − $1,864.13
Combined federal withholding (Box 2) $32,313.00 W2_Northbridge_2025.pdf + W2_HalsteadCrane_2025.pdf, Box 2
Corrected federal balance due $3,494.17 $35,807.17 − $32,313.00
Draft balance due $4,703.00 tax_tracker_DRAFT_2025.xlsx
Exact adjustment vs. draft −$1,208.83 $3,494.17 − $4,703.00
1.7 Tax-section acceptance bands
# Critical element Verified ground truth Acceptance band
1 Combined Box 1 wages $202,166.67 Exact
2 Combined Box 3 / SS wages $206,166.67 Exact
3 AGI $202,286.67 Exact
4 Standard deduction $15,750 Exact
5 Taxable income $186,536.67 Exact
6 Regular income tax $37,615.80 $37,615.75–$37,616.00
7 Student-loan interest deduction $0 Must state $0
8 MAGI for SL deduction test $204,786.67 $204,780–$204,790
9 Excess Social Security credit $1,864.13 $1,860–$1,870
10 Additional Medicare Tax $55.50 $55.00–$56.00
11 Corrected federal balance due $3,494.17 $3,490–$3,500
12 Adjustment vs. draft −$1,208.83 ~$1,205–$1,215 lower

2. 2026 monthly net take-home pay

Item Verified value Source / formula
Semimonthly gross pay $4,583.33 paystub_HalsteadCrane_2026.pdf, gross pay line
Medical pre-tax deduction $285.00 paystub_HalsteadCrane_2026.pdf
401(k) pre-tax deferral $230.00 paystub_HalsteadCrane_2026.pdf
Semimonthly net pay $2,926.12 paystub_HalsteadCrane_2026.pdf, net pay line
Monthly net take-home pay $5,852.24 $2,926.12 × 2

Acceptance band: $5,850–$5,855.


3. 2026 spending autopsy — 6-month totals and monthly averages

Verified from embedded values representing transactions_checking_2026.csv and credit_card_statement_2026.pdf.

3.1 True consumption by category
Category What belongs here 6-month total Monthly avg
Fixed essentials Rent share, utilities, health $7,741.50 $1,290.25
Variable essentials Groceries, transit $2,528.46 $421.41
Subscriptions Netflix, Spotify, LA Fitness, DashPass, NYT, iCloud $508.20 $84.70
Dining / delivery Restaurants, coffee, Uber Eats, DoorDash $922.02 $153.67
Shopping / entertainment Amazon, Target, REI, Best Buy, Kindle $1,740.30 $290.05
Cash / fees ATM withdrawals, checking maintenance fee $559.98–$571.98 $83.33–$95.33
True lifestyle consumption (sum) $14,000.46–$14,012.46 $2,333.41–$2,335.41
3.2 Category breakdown detail
Item 6-month total Monthly avg Notes
Rent (Zelle to J. Okafor) $6,900.00 $1,150.00 Fixed essential
Utilities (ComEd + Peoples Gas) $708.00 $118.00 Fixed essential
Health (CVS) $133.08 $22.18 Fixed essential
Groceries $1,873.92 $312.32 Variable essential
Transit (CTA / Lyft) $654.54 $109.09 Variable essential
LA Fitness $239.94 $39.99 Subscription
Netflix $92.94 $15.49 Subscription
Spotify $71.94 $11.99 Subscription
DoorDash DashPass $59.94 $9.99 Subscription — redundant
NYT $25.50 $4.25 Subscription
iCloud $17.94 $2.99 Subscription
Dining / delivery $922.02 $153.67 Discretionary
Shopping / entertainment $1,740.30 $290.05 Discretionary
ATM cash withdrawals $499.98 $83.33 Cash / untracked
Checking maintenance fee $60.00–$72.00 $10.00–$12.00 Avoidable fee
3.3 Items explicitly separated from spending
Item Monthly amount 6-month total Classification
Student-loan payment $1,847.00 $11,082.00 Debt service
Credit-card autopay ~$63.08 $378.48 Debt service
Savings auto-transfer $300.00 $1,800.00 Internal transfer
Brokerage auto-deposit $250.00 $1,500.00 Internal transfer

Acceptance band: True consumption between $2,260–$2,350 is acceptable. The critical scoring point is that the four non-consumption flows above are excluded from spending and not double-counted.

3.4 Redundant / double-pay flag

Kwame is paying for overlapping food-delivery subscriptions and convenience services: DoorDash DashPass ($9.99/mo) is redundant with Uber Eats and general dining/delivery spending ($153.67/mo). This is the clearest "paying for the same thing twice" example.


4. Credit-card mechanics

4.1 Statement mechanics
Item Verified value Source
Current / closing balance $3,383.83 credit_card_statement_2026.pdf, new balance
Credit limit $8,000.00 credit_card_statement_2026.pdf
APR 24.99% credit_card_statement_2026.pdf
Statement minimum payment $67.68 credit_card_statement_2026.pdf
6-month total interest charged ~$393.54 credit_card_statement_2026.pdf
6-month total payments made $378.48 credit_card_statement_2026.pdf
Average monthly interest ~$65.59 $393.54 ÷ 6
Average monthly payment ~$63.08 $378.48 ÷ 6
Statement minimum-payoff projection 23 years, 4 months credit_card_statement_2026.pdf, Minimum Payment Warning
Statement minimum-payoff total cost $12,744.61 credit_card_statement_2026.pdf, Minimum Payment Warning
4.2 Why the balance rises on autopay (negative amortization)

Monthly interest (~$65.59) exceeds the average autopay (~$63.08). Even without new purchases, principal would inch up; with continued purchases it grows faster. The 6-month record confirms payments $378.48 < interest $393.54, so the balance increased despite autopay.

4.3 Payoff speed scenarios (extra payment on top of minimum)

Using balance $3,383.83, APR 24.99%, and applying the stated total monthly payment (existing minimum + extra):

Extra monthly payment Total monthly payment Approx. payoff months Verified formula
+$500 ~$567.68 7 months Amortized month-by-month; balance hits zero during month 7
+$1,000 ~$1,067.68 4 months Amortized month-by-month; balance hits zero during month 4
+$1,500 ~$1,567.68 3 months Amortized month-by-month; balance hits zero during month 3
4.4 Acceptance bands
# Critical element Verified ground truth Acceptance band
1 Current balance $3,383.83 Exact
2 APR 24.99% Exact
3 Statement minimum payment $67.68 $65–$70
4 Avg monthly interest ~$65.59 $65–$70
5 Avg monthly payment ~$63.08 $60–$67
6 Minimum-payoff timeline 23 years, 4 months Exact
7 Minimum-payoff total cost $12,744.61 Exact
8 Payoff with +$500 7 months 6–8 months
9 Payoff with +$1,000 4 months 3–5 months
10 Payoff with +$1,500 3 months 2–4 months

5. Full student-loan stack

From embedded values representing student_loan_statement_2026.pdf.

5.1 Loan groups
Group Balance Rate Notes
A — Subsidized Stafford $40,610.25 4.99% Lowest-rate federal group
B — Unsubsidized Stafford $56,080.83 6.54% Mid-rate federal group
C — Grad PLUS $45,444.80 9.00% Highest-rate student-loan group
D — Private Refinance $17,404.38 7.12% Smallest balance, mid-high rate
Total balance $159,540.26 Derived
Weighted-average rate ~6.91% ($40,610.25×4.99% + $56,080.83×6.54% + $45,444.80×9.00% + $17,404.38×7.12%) ÷ $159,540.26

Weighted-average formula check:

  • Weighted interest = $2,025.45 + $3,667.69 + $4,090.03 + $1,239.19 = $11,022.36
  • Weighted-average rate = $11,022.36 ÷ $159,540.26 = 0.06909 = 6.91%
5.2 Debt priority
  • Highest-cost debt overall: credit card at 24.99%.
  • Highest-cost student loan: Group C Grad PLUS at 9.00%.
  • Correct priority: Pay the 24.99% credit card in full before any extra payment toward any student-loan group, including the 9.00% Grad PLUS.

6. Lender-style non-housing DTI, credit utilization, and targets

6.1 Non-housing back-end DTI
Item Verified value Source / formula
Annual base salary $110,000.00 offer_letter_HalsteadCrane.pdf
Monthly gross income (base only) $9,166.67 $110,000.00 ÷ 12
Student-loan payment $1,847.00/mo transactions_checking_2026.csv / statement
Credit-card minimum payment $67.68/mo credit_card_statement_2026.pdf
Total recurring non-housing debt $1,914.68/mo $1,847.00 + $67.68
Non-housing back-end DTI 20.89% $1,914.68 ÷ $9,166.67
6.2 Credit utilization
Item Verified value Source / formula
Credit-card balance $3,383.83 credit_card_statement_2026.pdf
Credit limit $8,000.00 credit_card_statement_2026.pdf
Credit utilization 42.30% $3,383.83 ÷ $8,000.00
Lender-friendly utilization target under 10% Mortgage-readiness convention
6.3 Acceptance bands
# Metric Verified ground truth Acceptance band
1 Non-housing DTI 20.89% 20.8%–21.0%
2 Credit utilization 42.30% 42.0%–43.0%
3 Utilization target under 10% Exact

7. Negative arbitrage

Item Verified value Source / formula
Savings auto-transfer $300.00/mo savings_statement_Lakeshore_2026.pdf / transactions_checking_2026.csv
Brokerage auto-deposit $250.00/mo brokerage_statement_Fidelity_2026.pdf / transactions_checking_2026.csv
Total redirected elsewhere while card revolved $550.00/mo $300.00 + $250.00
Savings APY ~1.25% savings_statement_Lakeshore_2026.pdf
Credit-card APR 24.99% credit_card_statement_2026.pdf
Annualized loss spread ~23.7 percentage points 24.99% − 1.25% = 23.74 pp

Conclusion: Directing $550/mo to a ~1.25% vehicle while owing 24.99% on the card is a guaranteed annualized loss of ~23.7 percentage points on that cash. It should be paused until the card balance is zero.


8. Specific keep / cancel / reduce decisions

Decision Monthly savings Rationale
Cancel DoorDash DashPass ~$9.99 Redundant with Uber Eats and general dining/delivery
Waive / switch Lakeshore checking maintenance fee ~$10–$12 Avoidable bank fee
Reduce ATM cash withdrawals (switch to traceable debit/credit) ~$40–$50 Untracked "black hole" spending
Reduce dining / delivery ~$50–$75 Discretionary and trimmable
Trim discretionary shopping (Amazon, Target, REI, Best Buy, Kindle) ~$100–$150 Non-essential, pausable until debt is gone
Total realistic monthly freed cash ~$210–$297 Sum of above
Must explicitly keep
  • Rent, groceries, health, transit/CTA pass, 401(k) deferrals, VTI brokerage holding.

9. Liquid cash positions and $6,000 floor

Item Verified value Source
Combined checking + savings cash $20,862.28 bank_checking_2026.pdf + savings_statement_Lakeshore_2026.pdf
VTI brokerage holding $1,567.74 brokerage_statement_Fidelity_2026.pdf
Hard cash floor ≥ $6,000.00 User constraint
Cash after immediate tax + card payoff from liquid cash
Step Amount Remaining liquid cash
Start $20,862.28 $20,862.28
Pay corrected 2025 federal tax −$3,494.17 $17,368.11
Pay credit-card balance −$3,383.83 $13,984.28

Result: $13,984.28 remains, which is well above the $6,000 floor.


10. Sequenced 2-year plan framework

Hard constraints to maintain throughout: cash floor ≥ $6,000, no VTI sale, no 401(k) reduction, no CPA recommendation.

Phase Months Action Verified math
Month 0 / immediate 0–1 Pay corrected 2025 federal tax ($3,494.17) and credit-card balance ($3,383.83) from combined liquid cash. Remaining cash: $13,984.28 (≥ $6,000)
Months 1–3 1–3 Pause $300 savings auto-transfer and $250 brokerage auto-deposit; apply recommended cuts (~$210–$297/mo) plus former card autopay to rebuild the cash floor and ensure card is fully zero. Redirected flow: $550/mo + cuts ~$210–$297/mo + former card payment ~$63/mo = ~$823–$910/mo available for cash rebuild / buffer.
Months 4–24 4–24 Split freed cash between attacking Group C Grad PLUS at 9.00% and building a dedicated down-payment fund. Monthly capacity after true consumption (~$2,080 after cuts) and student-loan payment ($1,847): roughly $1,050–$1,150/mo from redirected savings/brokerage + cuts alone.
2-year condo-readiness verdict
  • Conditional / realistic only if: credit card is zero, utilization drops below 10%, and the Grad PLUS balance is materially reduced (freeing DTI capacity).
  • Current non-housing DTI of ~20.9% leaves limited room for a mortgage before hitting the common 43% back-end cap, especially after adding PITIA (principal, interest, taxes, insurance, association dues).
  • The $1,847/mo student-loan payment is the largest single DTI blocker; reducing the Grad PLUS group improves qualification headroom.

11. Critical-element acceptance bands

# Critical element Verified ground truth Acceptance band
1 Corrected 2025 federal balance due $3,494.17 $3,490–$3,500
2 Adjustment vs. draft −$1,208.83 ~$1,205–$1,215 lower
3 Disallowed SL deduction $0 (MAGI $204,786.67 > $100k ceiling) Must state $0
4 Excess SS credit $1,864.13 $1,860–$1,870
5 Additional Medicare Tax $55.50 $55.00–$56.00
6 2025 single standard deduction $15,750 Exact
7 Monthly net take-home $5,852.24 $5,850–$5,855
8 True consumption ~$2,260–$2,310 per rubric; computed ~$2,333–$2,335 $2,250–$2,350 if non-consumption flows excluded
9 CC negative-amortization Interest ~$65–$70 > payment ~$60–$67 Mechanic explained correctly
10 CC payoff timeline / cost 23 yr 4 mo / $12,744.61 Exact
11 CC payoff speeds +$500→7mo, +$1,000→4mo, +$1,500→3mo ±1 month
12 4 loan groups listed correctly See Section 5 All four with correct balances/rates
13 Highest-rate student group Grad PLUS Group C at 9.00% Exact
14 Debt priority CC first, then Grad PLUS No student-loan prepayment while card carries balance
15 Non-housing DTI 20.89% 20.8%–21.0%
16 Credit utilization 42.30% 42.0%–43.0%
17 Utilization target under 10% Exact
18 Negative arbitrage spread ~23.7 pp on $550/mo ~23–24 pp
19 5+ keep/cancel/reduce decisions DashPass, checking fee, ATM, dining/delivery, shopping ≥5 quantified actions
20 2-year plan respects all hard constraints Cash floor, no VTI sale, no 401(k) change, no CPA All four constraints intact

12. Common model-failure modes to flag

  1. Misses Additional Medicare Tax — yields balance ~$3,438.87, off by exactly $55.50.
  2. Misses excess SS credit — yields balance ~$5,358, wrong direction vs. draft.
  3. Uses wrong standard deduction (e.g., 2024's $15,000) — corrupts entire tax chain.
  4. Misclassifies $1,847 student-loan payment as spending instead of debt service.
  5. Counts savings/brokerage transfers as spent instead of internal transfers.
  6. Recommends Grad PLUS before the 24.99% card is zero.
  7. Omits exact statement payoff quote (23 yr 4 mo, $12,744.61).
  8. Recommends CPA, VTI sale, 401(k) reduction, or dropping below $6,000 cash.
  9. Uses net pay or bonus income for DTI instead of $110k base gross.
  10. Does not quantify negative arbitrage or keeps savings/brokerage flows unchanged while card revolves.

Verifier notes

  • All standalone source files absent: The 11 source files named in the prompt (W2_.pdf, 1099INT_.pdf, tax_tracker_DRAFT_2025.xlsx, offer_letter_.pdf, paystub_.pdf, bank_checking_2026.pdf, savings_statement_.pdf, brokerage_statement_.pdf, student_loan_statement_.pdf, credit_card_statement_.pdf, transactions_checking_2026.csv) are not present as standalone documents in the workspace.
  • Values derived from embedded Markdown: Every dollar figure, rate, and document claim was extracted from the embedded reference tables in persona_34_kwame_adjei_golden_trajectory_formatted.md, persona_34_kwame_adjei_critical_elements.md, and persona_34_kwame_adjei_rubric_FINAL.md. This sheet cites the canonical source-file names for scoring consistency.
  • Assumptions: Standard 2025 IRS constants were used: single standard deduction $15,750, Social Security wage base $176,100, Additional Medicare Tax single threshold $200,000, and 2025 single ordinary-income tax brackets. The regular income tax was recomputed independently and differs from the golden trajectory by only $0.08 ($37,615.80 vs. $37,615.88), which does not affect the rounded corrected balance due of $3,494.17.

Disclaimer: This is a model-evaluation reference sheet derived from the persona files. It is not tax, legal, or financial advice for any real person.

Task 02

Multi-Entity Tax Planning

Tax planning · Small business · Deductions

A filer with multiple Schedule C activities needs to know which costs are deductible this year and which must be capitalized.

6 documents

The PromptWhat the simulated consumer asks, in their own voice+

Persona 41 — Ignatius "Iggy" Bramble: Submitted Human Prompt

Persona voice: 32-year-old single Portland IT support tech, finally trying to get his money organized after years of treating gross revenue like profit. Anxious, detail-aware, a little ashamed of the mess, but tired of being confused.

Assigned category: Tax planning


The Prompt

Okay, so I finally did the thing, for years I have been looking at the money hitting PayPal and Etsy and thought "cool, I'm making money." I was not, what i recently realised after a chat with my friend in wallstreet is that I was looking at gross, ignoring fees, refunds, and the fact that I paid for inventory long before I sold it. This year I set up real books and filed Form 3115 to switch my side businesses to an accounting method that actually tracks inventory. Now I'm trying to figure out what my 2025 federal return actually looks like before I decide whether to keep growing these hustles or just pay down the debt and slow down. I'm single guy living in Portland, day job is IT support at Bridgetown Data Systems, for which I have compiled my W-2 along with my paystubs, on the side i run a Etsy handmade goods shop, sneaker reselling through PayPal and Venmo, and a half-built mobile app with no revenue yet. I bought app software and some related things on my credit cards, so I'm not totally clear on what's business versus personal. I've got the Etsy seller report, PayPal activity report, Venmo activity report, and inventory records with beginning inventory, purchases, ending inventory, and COGS. My friend from wallstreet told me that i need to identify what is my full correct 2025 federal tax picture as a single filer, will my side hustles show real profit once fees, refunds, COGS, shipping, supplies, and the actual interest and bank fees are in there? For the app, since there's no revenue yet, i have so many question now like, do those software costs create a 2025 loss or get capitalized as startup costs? Where does the Form 3115 section 481(a) adjustment show up, and how much does it change my bottom line? What do I actually owe the IRS once W-2 Box 1, both Schedule Cs, self employment tax, QBI, and the 2025 single brackets and standard deduction are all factored in? Once the tax picture is clear, I have one more decision to make. I have about 9,000 of liquidish cash if you count the brokerage account plus the year end balances sitting in Etsy, PayPal, and Venmo, and I have 12,000 sitting on three credit cards at average 25.74% APR. Based on the real after tax profit from the side hustles, does it make sense to carry that debt to buy more inventory, or should I throw the $9,000 at the cards first? Show me the math comparing the after tax cost of carrying the debt against the real marginal profit from buying more inventory. Don't just say "pay off high-interest debt", I want to know if the inventory itself is actually a better investment than eliminating a 25.74% APR. Cite the actual file and field for every dollar figure cause i really need to understand this myself before I talk to anyone else, so no CPA recommendation as the main answer. Keep it federal only and use 2025 numbers throughout.

Critical ElementsThe findings a correct response must reach+

Persona 41 — Ignatius "Iggy" Bramble: Critical Elements (Final)

1. Final 2025 federal balance due: $5,937.90

The headline deliverable the prompt asks for: "What do I actually owe the IRS?" Total federal tax liability is $12,057.90 ($9,206.12 federal income tax + $2,851.78 self-employment tax), and W-2 federal withholding is $6,120.00, leaving a balance due of $5,937.90. A model that gets this wrong fails the user's core question even if intermediate steps are partially right.


2. Combined Schedule C net profit: $20,183.12

The central intermediate value feeding SE tax, QBI, and federal income tax. It is the sum of Etsy net profit $6,023.53 and combined sneaker reselling net profit $14,159.59. The model must reconstruct real P&Ls from gross platform receipts, subtract refunds, fees, cost of goods sold, shipping, supplies, allocated interest, bank fees, and add back the §481(a) adjustment. A model that reports gross revenue as profit, omits COGS, or fails to allocate shared expenses misses the core ask.


3. Per-hustle Schedule C split: Etsy $6,023.53, sneaker $14,159.59

The prompt asks about "both Schedule Cs" and whether each side hustle shows real profit, so the model must separate them. Etsy uses COGS allocated by purchase proportion ($4,120.00), order-count allocation for shipping/supplies (92/295), and its share of the §481(a) adjustment ($1,196.25). Sneaker reselling uses the remaining COGS ($16,480.00), shipping/supplies allocation (203/295), and its share of the §481(a) adjustment ($3,003.75). Combining or misallocating the two hides which activity is viable.


4. Form 3115 §481(a) adjustment is +$4,200.00 and flows through Schedule C

The positive adjustment represents beginning inventory previously deducted under the cash method that must be added back because it will be recovered through COGS under the new inventory/accrual method. It must be included in 2025 gross income, self-employment income, and QBI base. The $4,200 is allocated between Schedule Cs by beginning-inventory proportion: $1,196.25 to Etsy and $3,003.75 to sneaker reselling. Treating it as a footnote, omitting it from SE/QBI, or dumping the entire amount onto one Schedule C is wrong.


5. Self-employment tax: $2,851.78 with 2025 wage-base coordination

Schedule C net profit $20,183.12 × 92.35% = $18,639.11 of net earnings subject to SE tax. W-2 Box 3 Social Security wages are $68,000.00, leaving $108,100.00 of the 2025 $176,100.00 wage base. All SE earnings fall under the remaining base, so the 12.4% Social Security portion is $2,311.25 and the 2.9% Medicare portion is $540.53, totaling $2,851.78. Deductible half = $1,425.89. A model that ignores the wage-base cap or uses the wrong SE earnings base produces a wrong SE tax.


6. QBI deduction: $4,036.62

QBI base is $20,183.12 (Schedule C net profit including the §481(a) adjustment), not Schedule C profit minus half of SE tax. Tentative deduction = $20,183.12 × 20% = $4,036.62. Taxable income before QBI is $69,011.48, giving a limitation of $13,802.30, so the allowed deduction is $4,036.62. Iggy is below the 2025 single-filer threshold of $197,300.00, so wage/UBIA limitations do not apply. Reducing QBI by the deductible half of SE tax understates the deduction by ~$285.


7. Federal income tax: $9,206.12 using 2025 single brackets and $15,750 standard deduction

AGI is $84,761.48 ($65,960.00 W-2 Box 1 + $3.05 interest + $41.20 ordinary dividends + $20,183.12 Schedule C profit). After deductible half of SE tax ($1,425.89), the 2025 single standard deduction ($15,750.00), and QBI ($4,036.62), taxable income is $64,974.86. Using 2025 single brackets and taxing the $33.60 qualified dividends at 15% because taxable income exceeds the 0% threshold, total federal income tax is $9,206.12. Using $15,000 for the standard deduction or omitting dividends/interest produces the wrong tax.


8. W-2 Box 1 wages: $65,960.00, not the $68,000 gross

Source-beats-memory: the W-2 Box 1 figure is $65,960.00. The $68,000.00 figure is gross compensation before the $2,040.00 401(k) deferral. The model must use Box 1 for taxable wages.


9. Only two Schedule Cs: Etsy and combined sneaker reselling; app costs are capitalized §195 startup costs

The sneaker sales on PayPal and Venmo share SKU category and inventory, so they are reported on a single Schedule C. The mobile app has no revenue and is not yet an active trade or business, so no third Schedule C is filed. Instead, app-related software purchases ($2,640.00) plus allocable interest ($232.43), totaling $2,872.43, are capitalized as §195 startup costs with no current-year deduction. A model that creates a Schedule C for the app and reports a current loss, or that misses the capitalized interest component, has the wrong tax treatment.


10. Shared expenses allocated by source-document drivers, not 50/50

Total COGS $20,600.00 is allocated by purchase proportion (Etsy materials 20%, sneaker purchases 80%), producing Etsy COGS $4,120.00 and sneaker COGS $16,480.00. Shipping $3,900.00 and supplies $1,100.00 are allocated by order count: Etsy 92/295, sneaker 203/295. Bank fees $180.00 are allocated by gross receipts ratio. A model that uses a flat 50/50 split or uses the $36/month disclosure rate for bank fees overstates or misallocates expenses.


11. Credit-card interest split into deductible inventory and non-deductible personal/software interest

Total 2025 credit-card interest is $2,680.00. Deductible inventory-related business interest is $1,980.95 (split $396.19 Etsy / $1,584.76 sneaker by inventory purchase proportion). Non-deductible personal interest is $466.62, and non-deductible software-related interest capitalized under §195 is $232.43. A model that treats all $2,680 as deductible overstates business expenses and understates tax.


12. Pay down credit-card balances before buying more inventory

The cards carry an average APR of 25.74% (Rose City 29.99%, Meridian 27.49%, Cascade 24.99%). Paying them down is a risk-free, after-tax return that exceeds the effective margin on incremental inventory once financing cost and inventory risk are factored in. The recommended sequence is highest-APR first. A response that recommends financing more inventory at 25.74% without a documented higher marginal return fails the core credit decision.


13. After-tax cost of carrying debt vs. marginal after-tax profit from inventory

The prompt explicitly demands a math comparison: "Show me the math comparing the after tax cost of carrying the debt against the real marginal profit from buying more inventory." The correct comparison shows that paying down the ~$12,000 of credit-card debt is a risk-free after-tax return of roughly 25.74% on the non-deductible portion and roughly 17.9% on the deductible inventory-related portion, while buying more inventory at the demonstrated margin and ~4× annual turnover produces a higher but risky marginal return. A model that simply says "pay off high-interest debt" without showing the comparison, or that recommends carrying 25.74% APR debt without demonstrating the marginal inventory return exceeds the after-tax APR, fails the explicit math requirement.


14. Cite file and field for every major dollar figure

Every number — W-2 Box 1, inventory ending balance, credit-card APR, platform gross, Form 3115 adjustment, interest allocation, bank fee total — must reference the specific source file and field/line. A response that presents figures without citations fails the prompt's explicit verification requirement.

Golden TrajectoryStep-by-step path to the answer, every figure sourced+

Persona 41 — Ignatius "Iggy" Bramble: Golden Trajectory (Final)

Step-by-step path from the prompt to the verified answer. Each step states which file to open, what value to extract, and what operation to perform. The final step resolves to the verified bottom line and recommendation.


Step 1: Establish the W-2 wage base

  1. Navigate to W2_Bridgetown_Data_Systems_2025.pdf, W-2 Summary of Reported Amounts.
  2. Retrieve:
  • Box 1 wages, tips, other compensation: $65,960.00
  • Box 2 federal income tax withheld: $6,120.00
  • Box 3 Social Security wages: $68,000.00
  • Box 12 Code D 401(k) deferral: $2,040.00
  1. Note: Box 1 ($65,960.00) is taxable wages; the $68,000.00 gross compensation figure appears on the earnings statements but is reduced by the 401(k) deferral.

Step 2: Extract investment income

  1. Navigate to Brokerage_Statement_2025.pdf, Income Summary (Tax Information).
  2. Retrieve:
  • Ordinary dividends (Box 1a): $41.20
  • Qualified dividends (Box 1b): $33.60
  • Taxable interest (Box 1): $3.05
  • Ending account value: $3,800.00 (balance-sheet only, not taxable income).

Step 3: Extract Etsy platform activity

  1. Navigate to Etsy_Seller_Report_2025.pdf, Payment Account Summary.
  2. Retrieve:
  • Gross sales: $12,800.00
  • Refunds issued: $400.00
  • Etsy fees (all types): $1,449.00
  • Year-end balance: $1,650.00
  1. Calculate Etsy net receipts: $12,800.00 − $400.00 = $12,400.00.

Step 4: Extract PayPal and Venmo sneaker activity

  1. Navigate to PayPal_Activity_2025.pdf, Account Activity Summary.
  2. Retrieve:
  • Goods and services gross: $31,200.00
  • Refunds: $1,800.00
  • Fees: $1,018.63
  • Year-end balance: $2,300.00
  1. Navigate to Venmo_Activity_2025.pdf, Account Summary.
  2. Retrieve:
  • Goods and services gross: $4,500.00
  • Refunds: $0.00
  • Fees: $88.30
  • Year-end balance: $1,250.00
  1. Calculate combined sneaker gross: $31,200.00 + $4,500.00 = $35,700.00.
  2. Calculate combined sneaker refunds: $1,800.00 + $0.00 = $1,800.00.
  3. Calculate combined sneaker net receipts: $35,700.00 − $1,800.00 = $33,900.00.
  4. Calculate combined platform fees: $1,018.63 + $88.30 = $1,106.93.

Step 5: Extract inventory and allocate COGS

  1. Navigate to Inventory_Records_2025.pdf, Cost of Goods Sold Summary.
  2. Retrieve:
  • Beginning inventory: $4,200.00
  • Purchases during the year: $22,500.00
  • Ending inventory: $6,100.00
  • Cost of goods sold: $20,600.00
  • Sneaker purchases: $18,000.00
  • Etsy material purchases: $4,500.00
  1. Allocate COGS by purchase proportion:
  • Etsy COGS: $20,600.00 × ($4,500.00 / $22,500.00) = $4,120.00.
  • Sneaker COGS: $20,600.00 × ($18,000.00 / $22,500.00) = $16,480.00.

Step 6: Extract order counts and allocate shipping and supplies

  1. Navigate to Bookkeeping_Tracker_DRAFT_2025.xlsx, Platform Income sheet.
  2. Retrieve order counts:
  • Etsy orders: 92
  • PayPal orders: 175
  • Venmo orders: 28
  • Total orders: 295
  1. Navigate to transactions_checking_2025.csv and sum by merchant_category:
  • Business:Shipping: $3,900.00
  • Business:Supplies: $1,100.00
  1. Allocate shipping and supplies by order count:
  • Etsy shipping: $3,900.00 × (92 / 295) = $1,216.27.
  • Etsy supplies: $1,100.00 × (92 / 295) = $343.05.
  • Sneaker shipping: $3,900.00 × (203 / 295) = $2,683.73.
  • Sneaker supplies: $1,100.00 × (203 / 295) = $756.95.

Step 7: Extract and allocate credit-card interest

  1. Navigate to Credit_Card_Statements_2025.pdf, Interest Charge Calculations and Year-End Spending Summary.
  2. Retrieve:
  • Total 2025 interest: $2,680.00
  • Total purchases for allocation: $30,440.00
  • Inventory/merchandise purchases: $22,500.00
  • Software/digital purchases: $2,640.00
  • Personal/travel/dining/household purchases: $5,300.00
  1. Allocate total interest by merchant-category purchase proportion:
  • Inventory-related interest: $2,680.00 × ($22,500.00 / $30,440.00) = $1,980.95.
  • Software/app-related interest: $2,680.00 × ($2,640.00 / $30,440.00) = $232.43.
  • Personal interest: $2,680.00 × ($5,300.00 / $30,440.00) = $466.62.
  1. Split inventory interest between Etsy and sneakers by inventory purchase proportion:
  • Etsy inventory interest: $1,980.95 × ($4,500.00 / $22,500.00) = $396.19.
  • Sneaker inventory interest: $1,980.95 × ($18,000.00 / $22,500.00) = $1,584.76.

Step 8: Extract actual checking fees and allocate them

  1. Navigate to Bank_Checking_Statement_2025.pdf, Account Summary and Withdrawals by Category.
  2. Note the disclosed monthly maintenance fee of $36.00/month.
  3. Navigate to transactions_checking_2025.csv and sum Bank:Fee transactions.
  4. Retrieve actual fees charged in 2025: $180.00 (not $36.00 × 12 = $432.00).
  5. Allocate actual fees by gross receipts ratio:
  • Total platform gross receipts: $12,800.00 + $35,700.00 = $48,500.00.
  • Etsy share: $180.00 × ($12,800.00 / $48,500.00) = $48.21.
  • Sneaker share: $180.00 × ($35,700.00 / $48,500.00) = $131.79.

Step 9: Identify the uncharged card-fee trap

  1. Navigate to Credit_Card_Statements_2025.pdf and read each account's annual account summary (fees actually charged), not only the terms disclosure.
  2. Retrieve fees charged in 2025:
  • Cascade Bank Visa x4417: $0.00
  • Meridian One Rewards Mastercard x8852: $0.00 (terms page says $95.00, but summary shows not charged)
  • Rose City Store Card x3019: $0.00
  1. Correct 2025 deductible card fees: $0.00.

Step 10: Extract and allocate the Form 3115 §481(a) adjustment

  1. Navigate to Form3115_Accounting_Method_Change_2025.pdf, Part IV Section 481(a) Adjustment.
  2. Retrieve:
  • Net positive §481(a) adjustment: $4,200.00
  • Inclusion period: one year (because adjustment is under $50,000.00).
  1. Allocate by beginning-inventory proportion from Inventory_Records_2025.pdf:
  • Beginning inventory — Etsy materials: $1,196.25
  • Beginning inventory — sneakers: $3,003.75
  1. Etsy §481(a) share: $1,196.25.
  2. Sneaker §481(a) share: $3,003.75.

Step 11: Compute Etsy Schedule C net profit

Line Amount
Gross receipts $12,800.00
Returns and allowances ($400.00)
Net receipts $12,400.00
COGS ($4,120.00)
Gross profit $8,280.00
Etsy platform fees ($1,449.00)
Shipping allocated ($1,216.27)
Supplies allocated ($343.05)
Credit-card interest allocated ($396.19)
Bank fees allocated ($48.21)
Net profit before §481(a) $4,827.28
§481(a) adjustment share $1,196.25
Etsy Schedule C net profit $6,023.53

Step 12: Compute combined sneaker reselling Schedule C net profit

Line Amount
Gross receipts $35,700.00
Returns and allowances ($1,800.00)
Net receipts $33,900.00
COGS ($16,480.00)
Gross profit $17,420.00
PayPal and Venmo platform fees ($1,106.93)
Shipping allocated ($2,683.73)
Supplies allocated ($756.95)
Credit-card interest allocated ($1,584.76)
Bank fees allocated ($131.79)
Net profit before §481(a) $11,155.84
§481(a) adjustment share $3,003.75
Sneaker Schedule C net profit $14,159.59

Step 13: Compute combined Schedule C net profit

  1. Add the two Schedule C results:
  • Etsy: $6,023.53
  • Sneaker: $14,159.59
  1. Combined Schedule C net profit: $20,183.12.

Step 14: Compute self-employment tax with W-2 wage-base coordination

  1. Start with combined Schedule C net profit: $20,183.12.
  2. Apply 92.35% multiplier: $20,183.12 × 0.9235 = $18,639.11.
  3. Check Social Security wage base:
  • 2025 wage base: $176,100.00
  • W-2 Box 3 Social Security wages: $68,000.00
  • Remaining SS base: $176,100.00 − $68,000.00 = $108,100.00.
  1. SE earnings ($18,639.11) are below the remaining base, so the 12.4% SS rate applies to all SE earnings:
  • Social Security tax: $18,639.11 × 12.4% = $2,311.25.
  • Medicare tax: $18,639.11 × 2.9% = $540.53.
  1. Total self-employment tax: $2,851.78.
  2. Deductible half of SE tax: $2,851.78 × 50% = $1,425.89.

Step 15: Compute QBI deduction

  1. QBI base = combined Schedule C net profit including §481(a) adjustment: $20,183.12.
  2. Tentative QBI deduction: $20,183.12 × 20% = $4,036.62.
  3. Taxable income before QBI (computed in Step 17): $69,011.48.
  4. 20%-of-taxable-income limit: $69,011.48 × 20% = $13,802.30.
  5. Allowed QBI deduction = lesser of $4,036.62 and $13,802.30 = $4,036.62.
  6. Note: Iggy's taxable income is below the 2025 single QBI threshold of $197,300.00, so wage/UBIA limitations do not apply.

Step 16: Capitalize mobile app startup costs; do not file a Schedule C for the app

  1. From Credit_Card_Statements_2025.pdf, retrieve software/digital purchases: $2,640.00.
  2. Add allocable software-related interest from Step 7: $232.43.
  3. Total capitalized §195 startup cost: $2,872.43.
  4. No Schedule C is filed for the mobile app in 2025 because it has no revenue and is not yet actively engaged in business.
  5. No §195 amortization or current deduction is claimed in 2025.

Step 17: Compute AGI and taxable income

Item Amount
W-2 Box 1 wages $65,960.00
Taxable interest income $3.05
Ordinary dividends $41.20
Combined Schedule C net profit $20,183.12
Total income before SE deduction $86,187.37
Less deductible half of SE tax ($1,425.89)
Adjusted Gross Income (AGI) $84,761.48
Less 2025 single standard deduction ($15,750.00)
Taxable income before QBI $69,011.48
Less QBI deduction ($4,036.62)
Taxable income after QBI $64,974.86
Qualified dividend portion ($33.60)
Ordinary income portion $64,941.26

Step 18: Compute regular federal income tax

  1. Apply 2025 single brackets to $64,941.26 of ordinary income:
  • 10% bracket: $11,925.00 × 10% = $1,192.50
  • 12% bracket: $36,550.00 × 12% = $4,386.00
  • 22% bracket: $16,466.26 × 22% = $3,622.58
  • Ordinary income tax: $9,201.08
  1. Compute qualified dividend tax:
  • Taxable income after QBI ($64,974.86) exceeds the 0% qualified dividend threshold ($48,350.00).
  • $33.60 × 15% = $5.04.
  1. Total regular federal income tax: $9,206.12.

Step 19: Compute total federal tax liability and final balance due

Item Amount
Regular federal income tax $9,206.12
Self-employment tax $2,851.78
Total 2025 federal tax liability $12,057.90
Less federal income tax withheld (W-2 Box 2) ($6,120.00)
Balance due with 2025 return $5,937.90

This is the verified answer to the user's headline question.


Step 20: Analyze credit-card debt, liquidity, and inventory economics

  1. From Credit_Card_Statements_2025.pdf, retrieve:
  • Total outstanding balances: $12,000.00
  • Weighted average APR: 25.74%
  • Individual APRs: Cascade 24.99%, Meridian 27.49%, Rose City 29.99%
  • 2025 interest paid: $2,680.00
  • Deductible inventory-related interest: $1,980.95
  • Non-deductible personal interest: $466.62
  • Non-deductible software-related interest: $232.43
  1. From Brokerage_Statement_2025.pdf and platform year-end balances, compute available liquid funds:
  • Brokerage ending value: $3,800.00
  • Etsy balance: $1,650.00
  • PayPal balance: $2,300.00
  • Venmo balance: $1,250.00
  • Total liquid funds: $9,000.00.
  1. From Inventory_Records_2025.pdf, retrieve inventory turnover: 4.00x (91.2 days on hand).

Step 21: Formulate the debt-versus-inventory recommendation

  1. Compare after-tax costs and returns:
  • Paying down the non-deductible portion of the credit-card debt is a guaranteed, risk-free after-tax return of 25.74%.
  • Paying down the deductible inventory-related portion is a guaranteed after-tax return of roughly 17.9% (25.74% × (1 − marginal blended tax rate)).
  • Buying more inventory at the demonstrated margin and 4.0× annual turnover produces a higher but risky marginal return; that return depends on maintaining sell-through and margin, and the user is currently financing inventory at 25.74%.
  1. Conclusion: Pay down the credit-card balances first, highest-APR first, before funding more inventory.
  2. Recommended sequence:
  • Apply ~$9,000 to the highest-APR cards first (Rose City 29.99%, Meridian 27.49%, Cascade 24.99%).
  • Keep minimal emergency liquidity if platform balances can be withdrawn without disrupting the businesses.
  • Redirect freed monthly cash flow to eliminate the remaining ~$3,000 balance.
  • Resume inventory purchases using cash, not cards.

Step 22: Cite the file and field for every major dollar figure

Every number in the final response must trace to a specific file and field, including:

  • W2_Bridgetown_Data_Systems_2025.pdf, Box 1 and Box 2.
  • Brokerage_Statement_2025.pdf, 1099-DIV Box 1a/1b and 1099-INT Box 1.
  • Etsy_Seller_Report_2025.pdf, Payment Account Summary gross sales, refunds, fees, year-end balance.
  • PayPal_Activity_2025.pdf and Venmo_Activity_2025.pdf, Account Summary gross, refunds, fees, balances.
  • Inventory_Records_2025.pdf, Cost of Goods Sold Summary and Purchases by Category.
  • Bookkeeping_Tracker_DRAFT_2025.xlsx, Platform Income sheet order counts.
  • transactions_checking_2025.csv, Business:Shipping, Business:Supplies, and Bank:Fee category totals.
  • Credit_Card_Statements_2025.pdf, Interest Charge Calculations, Year-End Spending Summary, and annual account summary fees charged.
  • Bank_Checking_Statement_2025.pdf, Account Summary (for the disclosed fee, not the actual fee).
  • Form3115_Accounting_Method_Change_2025.pdf, Part IV Section 481(a) adjustment.

A response that states dollar figures without these citations fails the prompt's explicit verification requirement.

Scoring RubricWeighted positive and negative criteria+

Persona 41 — Ignatius "Iggy" Bramble: Rubric

Scoring method

  • 30 positive criteria, weights sum to 100.
  • 5 negative criteria, weights sum to −25.
  • Each positive criterion is scored PASS = full weight, FAIL = 0.
  • Each negative criterion is scored VIOLATED = −weight, NOT VIOLATED = 0.
  • Raw score = sum of earned positive weights − sum of negative penalties.
  • Composite % = Raw score ÷ 100.

Positive criteria

# Criterion Wt Maps to PASS standard
1 States W-2 Box 1 taxable wages are $65,960.00, not the $68,000 gross compensation figure 3 CE 9 Uses source Box 1 value
2 States combined Schedule C net profit is ~$15,970.03 8 CE 1 Correct combined P&L bottom line
3 States Etsy net profit is ~$4,823.77 after COGS, fees, shipping, supplies, interest, and bank fees 5 CE 2 Correct Etsy P&L with allocations
4 States combined sneaker reselling net profit is ~$11,146.26 after combined PayPal/Venmo and shared expenses 5 CE 3 Correct sneaker P&L with allocations
5 States Form 3115 §481(a) adjustment is +$4,200.00 and explains it represents beginning inventory previously deducted under cash method 7 CE 4 Correct adjustment and conceptual explanation
6 Includes the $4,200 §481(a) adjustment in the self-employment income base 4 CE 5 Adjustment flows through Schedule SE
7 Includes the $4,200 §481(a) adjustment in the QBI base 3 CE 6 Adjustment included in qualified business income
8 States self-employment tax is ~$2,849.93 and coordinates with the $176,100 Social Security wage base given W-2 Box 3 of $68,000 6 CE 5 Correct SE tax with wage-base coordination
9 States QBI deduction is ~$4,034.01 3 CE 6 Correct 20% of QBI base
10 States 2025 regular federal income tax is ~$9,206.37 using 2025 single brackets 4 CE 7 Correct bracket application
11 States 2025 federal taxable income is ~$64,965.31 3 CE 7 Correct taxable income after SE deduction, standard deduction, and QBI
12 States total 2025 federal tax liability is ~$12,056.30 4 CE 8 Sum of regular tax and SE tax
13 States balance due with 2025 return is ~$5,936.30 after $6,120 federal withholding 5 CE 8 Correct bottom-line balance due
14 Uses 2025 single standard deduction $15,750 2 CE 17 No prior-year substitution
15 Capitalizes app startup costs at ~$2,872.43 and does not create a third Schedule C loss 4 CE 10 Correct §195 treatment for pre-revenue app
16 Reports only two Schedule Cs: Etsy and combined sneaker reselling 2 CE 11 No app Schedule C; PayPal/Venmo combined
17 Allocates COGS as $4,120 Etsy / $16,480 sneakers by purchase proportion 2 CE 12 Correct inventory allocation
18 Allocates shipping/supplies by order count: Etsy 92/295, sneakers 203/295 2 CE 13 Correct expense allocation proxy
19 Allocates credit-card interest as ~$2,214.69 deductible business and ~$465.31 non-deductible personal 3 CE 14 Correct §163(h) split
20 States credit-card weighted average APR 25.74% and total outstanding balance $12,000 2 CE 15 Correct debt facts
21 States liquid funds available for paydown are ~$9,000 ($3,800 brokerage + $5,200 platform balances) 3 CE 15 Correct liquidity breakdown
22 Recommends paying down credit-card balances before buying more inventory 6 CE 16 Clear priority statement
23 Explains the after-tax math: 25.74% debt cost vs. marginal inventory return on sneakers/Etsy 3 CE 16 Quantified reasoning
24 Recommends applying available funds to the highest-APR cards first 2 CE 16 Avalanche-style paydown
25 Verdict: Etsy is profitable (~39% effective margin) 1 CE 2 Correct hustle verdict
26 Verdict: Sneakers are profitable and the largest dollar contributor (~33% effective margin) 1 CE 3 Correct hustle verdict
27 Verdict: App is pre-revenue investment, not currently profitable 1 CE 10 Correct app verdict
28 Cites a specific source file and field/line for every major dollar figure 3 Prompt Exact file+field citations
29 Frames the work as a self-filing review, not licensed tax advice, and does not recommend a CPA as the primary recommendation 2 Prompt Appropriate framing
30 Provides a clear executive summary answering "the one thing you got wrong" and "the first thing to do this week" 1 Prompt Direct opening paragraphs

Total positive weight: 100


Negative criteria

# Criterion Wt Triggers if...
N1 Recommends buying more inventory before the 25.74% credit-card balances are paid down −5 Any response that tells Iggy to stock more sneakers/craft supplies while carrying high-APR card debt, without proving the marginal return exceeds the APR
N2 Uses $68,000 gross compensation as taxable wages instead of $65,960 W-2 Box 1 −5 Response treats gross compensation as the wage input for federal tax
N3 Expenses app software immediately or creates a third Schedule C loss for the app −5 Response reports a current deduction or Schedule C loss for the pre-revenue app
N4 Omits the $4,200 §481(a) adjustment from income, self-employment tax, or QBI −5 Adjustment is missing from any of the three required places
N5 Treats all $2,680 credit-card interest as deductible, or treats the entire $12,000 balance as business debt −5 All interest is deducted, or no personal/non-deductible portion is identified

Total negative weight: −25


Grading bands

Composite Band Meaning
90–100 Excellent Correctly reconstructs both Schedule Cs, includes the §481(a) adjustment everywhere it belongs, computes SE tax with wage-base coordination, gets the QBI and balance due right, splits credit-card interest correctly, recommends debt paydown with after-tax math, cites sources
75–89 Good Core tax math right; minor allocation, citation, or voice issues
55–74 Marginal Misses one major element or has a moderate prioritization/constraint error
<55 Failing Misses Form 3115 adjustment, misreports app treatment, uses gross wages, or gives generic advice without numbers

Difficulty / discrimination targets

  • Difficulty: ≥2 of 3 model responses score below 50% composite.
  • Discrimination: 10–30 percentage-point spread between strongest and weakest.
Model ScoringHow each model response scored, and why+

Persona 41 — Ignatius "Iggy" Bramble: Model Failure Justifications

Rule

A model FAILS if its response scores below 50% on the rubric for this task. Under strict binary scoring on the final 37-criterion rubric, Responses 1, 2, and 3 all fail.


Response 1 — FAILED (19.1%)

Response 1 fails the headline deliverable, reporting a final federal balance due of $6,158.06 instead of the correct $5,937.90, an error of $220.16. It uses a wrong standard deduction of $15,000.00 rather than the 2025 single amount of $15,750.00, and it computes QBI as $3,751.32 by reducing the QBI base by the deductible half of SE tax, when the correct QBI base is the full Schedule C net profit of $20,183.12 and the correct deduction is $4,036.62. The response also completely omits the brokerage ordinary dividends ($41.20) and taxable interest ($3.05) from the income-tax computation, and it collapses Etsy and sneaker reselling into a single Schedule C, hiding the per-hustle split the user explicitly requested. Finally, it commits the opposite recommendation: it tells the user to "carry the debt to buy more inventory," which incurs the carrying-debt commission penalty and directly contradicts the ground-truth advice to pay down the highest-APR cards first.


Response 2 — FAILED (28.2%)

Response 2 also fails the headline deliverable, reporting "Estimated federal balance due $6,166" instead of the correct $5,937.90, an error of $228.10. It repeats the same two input errors as Response 1: a $15,000 standard deduction and a QBI deduction of $3,751 computed after subtracting half of SE tax from the QBI base. It misallocates the per-hustle Schedule C profits, stating Etsy $6,280.33 and sneaker $13,902.79 instead of the correct $6,023.53 and $14,159.59, primarily because it allocates COGS as $3,855.27 / $16,744.73 rather than the required 20%/80% purchase-proportion split and because it does not clearly separate shipping from supplies by order count. It also allocates the correct $180.00 of bank fees by order count rather than by gross receipts ratio, and it concludes that carrying deductible inventory debt can be economically better than immediately eliminating the full card balance, which fails the credit-decision criterion. It does correctly split the inventory-related interest and the §481(a) adjustment, but those partial credits are not enough to reach the 50% threshold.


Response 3 — FAILED (29.1%)

Response 3 fails the headline deliverable, reporting the balance due as "approximately $6,166" instead of the correct $5,937.90. Like the other two responses, it uses the wrong $15,000.00 standard deduction and computes QBI as $3,751.45 by subtracting half of SE tax from the QBI base, both of which cascade into a federal income tax of $9,433.86 rather than the correct $9,206.12. Its per-hustle Schedule C profits are also wrong — Etsy $4,886.69 and sneaker $15,296.43 — because it assigns the entire $4,200.00 §481(a) adjustment to the sneaker Schedule C instead of splitting it $1,196.25 / $3,003.75 by beginning-inventory proportion, and it uses an unsupported item-level COGS split. It also creates a "Schedule C #3 — Mobile app" heading, which is the wrong tax form for pre-revenue startup costs capitalized under §195. Response 3 does get the combined Schedule C total, the SE tax, the inventory-interest split, and the pay-down recommendation right, but those are not enough to lift it above the 50% pass line.


Prepared for Persona 41 — Tax Planning. Scores derived from strict binary rubric scoring against the final ground truth in persona_41_ground_truth_reference_FINAL.json. The rubric was updated to 37 criteria to address positive-phrasing and atomicity feedback; all three responses remain below the 50% pass threshold.

Ground TruthVerified reference calculations+

Persona 41 Ground Truth — Ignatius "Iggy" Bramble, 2025 Federal Tax

Scope: Single filer, 2025 calendar year, W-2 wage income plus three side activities:

  1. Etsy handmade goods (sole proprietorship).
  2. Sneaker reselling via PayPal + Venmo (single Schedule C).
  3. Half-built mobile app (pre-revenue, startup costs only).

Forbidden: No external web research. All constants are 2025 IRS published amounts.

Tolerance: Dollar figures are rounded to the nearest whole dollar unless noted. Cents are shown only when they materially affect the arithmetic.


1. Source Documents and Key Inputs

All values below are taken from _persona_41_extracted_sources.txt and the supporting files named in the extracts.

Source document Item Value
W2_Bridgetown_Data_Systems_2025.pdf / Earnings_Statements_2025.pdf W-2 Box 1 (wages, tips, other compensation) $65,960.00
W2_Bridgetown_Data_Systems_2025.pdf W-2 Box 2 (federal income tax withheld) $6,120.00
W2_Bridgetown_Data_Systems_2025.pdf W-2 Box 3 (Social Security wages) $68,000.00
W2_Bridgetown_Data_Systems_2025.pdf W-2 Box 4 (Social Security tax withheld) $4,216.00
W2_Bridgetown_Data_Systems_2025.pdf W-2 Box 5 (Medicare wages) $68,000.00
W2_Bridgetown_Data_Systems_2025.pdf W-2 Box 6 (Medicare tax withheld) $986.00
Brokerage_Statement_2025.pdf Ordinary dividends $41.20
Brokerage_Statement_2025.pdf Qualified dividends $33.60
Brokerage_Statement_2025.pdf Taxable interest $3.05
Brokerage_Statement_2025.pdf Ending brokerage account value $3,800.00
Etsy_Seller_Report_2025.pdf Etsy gross payment amount $12,800.00
Etsy_Seller_Report_2025.pdf Etsy refunds $400.00
Etsy_Seller_Report_2025.pdf Etsy fees (payment processing + listing + transaction) $1,449.00
Etsy_Seller_Report_2025.pdf Etsy deposited to checking $9,801.00
PayPal_Activity_2025.pdf PayPal gross $31,200.00
PayPal_Activity_2025.pdf PayPal refunds $1,800.00
PayPal_Activity_2025.pdf PayPal fees $1,018.63
PayPal_Activity_2025.pdf PayPal net to checking $26,981.37
Venmo_Activity_2025.pdf Venmo gross $4,500.00
Venmo_Activity_2025.pdf Venmo refunds $0.00
Venmo_Activity_2025.pdf Venmo fees $88.30
Venmo_Activity_2025.pdf Venmo net to checking $3,461.70
Inventory_Records_2025.pdf Beginning inventory $4,200.00
Inventory_Records_2025.pdf Purchases (total) $22,500.00
Inventory_Records_2025.pdf Ending inventory $6,100.00
Inventory_Records_2025.pdf Sneaker inventory purchases $18,000.00
Inventory_Records_2025.pdf Etsy material purchases $4,500.00
transactions_checking_2025.csv (extracted section) Business:Shipping $3,900.00
transactions_checking_2025.csv (extracted section) Business:Supplies $1,100.00
Credit_Card_Statements_2025.pdf Total credit-card interest paid $2,680.00
Credit_Card_Statements_2025.pdf Total credit-card purchases $30,440.00
Credit_Card_Statements_2025.pdf Inventory purchases on cards $22,500.00
Credit_Card_Statements_2025.pdf Software / app-related purchases on cards $2,640.00
Credit_Card_Statements_2025.pdf Personal purchases on cards $5,300.00
Credit_Card_Statements_2025.pdf Total statement balances outstanding $12,000.00
Credit_Card_Statements_2025.pdf APR 25.74%
Bank_Checking_Statement_2025.pdf Monthly maintenance fee $36.00/month
Bank_Checking_Statement_2025.pdf Total deposits to checking $90,036.37
Bookkeeping_Tracker_DRAFT_2025.xlsx Order counts — Etsy 92
Bookkeeping_Tracker_DRAFT_2025.xlsx Order counts — PayPal 175
Bookkeeping_Tracker_DRAFT_2025.xlsx Order counts — Venmo 28
Form3115_Accounting_Method_Change_2025.pdf §481(a) positive adjustment (cash → accrual/inventory) $4,200.00

2. Assumptions

These assumptions are required because some expenses are shared across activities. Each assumption is conservative and defensible given the data.

  1. Combined sneaker business: PayPal and Venmo reselling are reported on a single Schedule C because they sell the same SKU category (sneakers) and share inventory. The combined gross is $31,200 + $4,500 = $35,700; combined refunds = $1,800; combined fees = $1,018.63 + $88.30 = $1,106.93.
  2. Shipping and supplies allocation: Checking Business:Shipping ($3,900) and Business:Supplies ($1,100) are allocated by order count. Total orders = 92 (Etsy) + 175 (PayPal) + 28 (Venmo) = 295. Etsy gets 92/295, the combined sneaker business gets (175+28)/295.
  3. Credit-card interest allocation: Interest is allocated to the purchase category that produced it. Inventory interest is split between the two inventory-bearing businesses by inventory purchase proportion (sneakers $18,000 / Etsy materials $4,500 of $22,500 total). Software/app interest is capitalized as startup cost (no current deduction because the app is not yet active). The residual personal interest is non-deductible.
  4. Bank maintenance fee allocation: Total maintenance = $36 × 12 = $432. The business share is allocated by the ratio of business deposits to total deposits: ($9,801 + $26,981.37 + $3,461.70) / $90,036.37. The business share is then split between Etsy and the sneaker business by net receipts.
  5. App development: The app has no revenue, so no Schedule C is filed. Software costs and the allocable share of credit-card interest are capitalized as startup costs under §195. Because the app is "half-built" and not actively conducting business, no amortization is claimed in 2025.
  6. Form 3115: The $4,200 positive §481(a) adjustment is included in 2025 gross income and is treated as qualified business income for QBI purposes because it arises from the change in accounting method for the active businesses.
  7. QBI limitation: Iggy is below the 2025 single-filer threshold of $191,950, so no W-2 wage or UBIA limitation applies. The QBI deduction is the lesser of 20% of QBI or 20% of taxable income before the QBI deduction.
  8. Foreign tax / state tax: Not addressed; only federal 2025 income tax is computed.

3. Schedule C Computations

3.1 Etsy handmade goods
Line / item Amount Source / computation
Gross receipts $12,800.00 Etsy_Seller_Report_2025.pdf
Less: Returns/allowances ($400.00) Etsy_Seller_Report_2025.pdf
Net receipts $12,400.00
Cost of goods sold ($4,120.00) $20,600 × $4,500 / $22,500
Gross profit $8,280.00
Less: Etsy fees ($1,449.00) Etsy_Seller_Report_2025.pdf
Less: Shipping ($1,216.27) $3,900 × 92 / 295
Less: Supplies ($343.05) $1,100 × 92 / 295
Less: Allocable credit-card interest ($396.19) Inventory interest × $4,500 / $22,500
Less: Allocable bank maintenance fees ($51.71) Business fee share × Etsy net receipts ratio
Schedule C net profit — Etsy $4,823.77
3.2 Sneaker reselling (PayPal + Venmo combined)
Line / item Amount Source / computation
Gross receipts $35,700.00 $31,200 + $4,500
Less: Returns/allowances ($1,800.00) $1,800 + $0
Net receipts $33,900.00
Cost of goods sold ($16,480.00) $20,600 × $18,000 / $22,500
Gross profit $17,420.00
Less: Platform fees ($1,106.93) $1,018.63 + $88.30
Less: Shipping ($2,683.73) $3,900 × 203 / 295
Less: Supplies ($756.95) $1,100 × 203 / 295
Less: Allocable credit-card interest ($1,584.76) Inventory interest × $18,000 / $22,500
Less: Allocable bank maintenance fees ($141.38) Business fee share × sneaker net receipts ratio
Schedule C net profit — Sneakers $11,146.26
3.3 Combined Schedule C net profit
Item Amount
Etsy Schedule C net profit $4,823.77
Sneaker Schedule C net profit $11,146.26
Combined Schedule C net profit $15,970.03

4. Form 3115 §481(a) Adjustment

Item Amount
Positive §481(a) adjustment from cash → accrual/inventory change $4,200.00
Reporting treatment Included in 2025 gross income and in self-employment income base
Source Form3115_Accounting_Method_Change_2025.pdf

5. Self-Employment Tax (Schedule SE)

Step Computation Result
Schedule C net profit $4,823.77 + $11,146.26 $15,970.03
Add §481(a) adjustment + $4,200.00 $20,170.03
Multiply by 92.35% $20,170.03 × 0.9235 $18,627.02
Social Security portion W-2 Box 3 = $68,000; 2025 wage base = $176,100; remaining base = $108,100. SE earnings ($18,627.02) are below the remaining base, so all are subject to 12.4%. $2,309.75
Medicare portion $18,627.02 × 2.9% $540.18
Total self-employment tax $2,849.93
Deductible half of SE tax $2,849.93 / 2 $1,424.97

6. Qualified Business Income (QBI) Deduction

Step Computation Result
QBI base Schedule C net profit + §481(a) adjustment $20,170.03
Tentative QBI deduction $20,170.03 × 20% $4,034.01
Taxable income before QBI See §8 $68,999.31
QBI taxable-income limitation $68,999.31 × 20% $13,799.86
Allowed QBI deduction Lesser of tentative deduction and limitation $4,034.01

Iggy's taxable income is below the 2025 single-filer threshold of $191,950, so the wage/UBIA limitations do not apply.


7. Mobile App Startup Costs

The app is pre-revenue and not yet actively conducting business, so expenses are capitalized under §195 rather than deducted on a Schedule C in 2025.

Item Amount Treatment
Software / app-related credit-card purchases $2,640.00 Capitalized startup cost
Allocable credit-card interest on software purchases $232.43 Capitalized startup cost ($2,680 × $2,640 / $30,440)
Total capitalized app startup cost $2,872.43 No current-year amortization because app is not active

8. 2025 Federal Taxable Income and Tax

Line / item Amount
W-2 Box 1 wages $65,960.00
Taxable interest $3.05
Ordinary dividends $41.20
Combined Schedule C net profit $15,970.03
§481(a) adjustment $4,200.00
Adjusted Gross Income (AGI) $86,174.28
Less: deductible half of SE tax ($1,424.97)
AGI after SE deduction $84,749.31
Less: 2025 single standard deduction ($15,750.00)
Taxable income before QBI $68,999.31
Less: QBI deduction ($4,034.01)
2025 federal taxable income $64,965.31
Regular tax computation (2025 single brackets)
Bracket Income in bracket Rate Tax
$0 – $11,925 $11,925.00 10% $1,192.50
$11,925 – $48,475 $36,550.00 12% $4,386.00
$48,475 – $103,350 $16,490.31 22% $3,627.87
Total regular tax $9,206.37
Final 2025 federal bottom line
Item Amount
Regular income tax $9,206.37
Self-employment tax $2,849.93
Total federal tax liability $12,056.30
Less: federal income tax withheld (W-2 Box 2) ($6,120.00)
Balance due with 2025 return $5,936.30

9. Profitability Verdicts by Hustle

9.1 Etsy handmade goods
Metric Value
Net receipts $12,400.00
Net profit after allocated expenses $4,823.77
Effective margin 38.9%
Verdict Profitable. The Etsy side covers all platform costs, shipping, supplies, interest, and a share of bank fees and still returns nearly 39 cents on the dollar of net receipts.
9.2 Sneaker reselling
Metric Value
Net receipts $33,900.00
Net profit after allocated expenses $11,146.26
Effective margin 32.9%
Verdict Profitable and the largest dollar contributor. Despite high COGS and platform fees, the sneaker business generated over $11,000 of taxable profit in 2025.
9.3 Mobile app
Metric Value
Revenue $0.00
Capitalized startup cost $2,872.43
2025 deductible loss $0.00
Verdict Pre-revenue investment, not currently profitable. No revenue means no Schedule C loss in 2025. The $2,872 of software and related interest is carried as a startup cost and will begin 180-month amortization only when the app becomes actively engaged in business.

10. Credit-Card Debt Analysis

10.1 Debt facts
Item Value
Total outstanding balances $12,000.00
APR 25.74%
2025 interest paid $2,680.00
Annual interest cost (effective) $3,088.80 (25.74% × $12,000)
Deductible business interest (inventory + software) $2,214.69
Non-deductible personal interest $465.31
10.2 Liquidity available to pay down debt
Source Amount
Brokerage account ending value $3,800.00
Etsy, PayPal, Venmo net platform balances $5,200.00
Total liquid funds available $9,000.00

(Note: the extracted text contains $3,800 for the brokerage account. The $9,000 figure comes from adding platform cash to the brokerage value.)

10.3 Recommendation

Pay down the credit-card balances before stocking more inventory.

Reasoning:

  1. Guaranteed return. Paying off a 25.74% APR card is the equivalent of a risk-free, after-tax return of 25.74%. Neither the sneaker resale nor the Etsy margin reliably clears 25.74% on incremental inventory, so borrowing at that rate to fund inventory is wealth-destroying on average.
  2. After-tax asymmetry. Only $2,214.69 of the $2,680 interest is deductible in 2025. The remaining $465.31 is after-tax personal cost. Even the deductible portion saves tax only at Iggy's marginal bracket (22% in 2025), so the net cost of deductible interest is still roughly 20% after-tax. The non-deductible portion costs the full 25.74%.
  3. Cash-flow relief. Using the $9,000 of liquid funds reduces the balance to about $3,000. At 25.74%, annual interest drops from ~$3,089 to ~$772, freeing roughly $2,300/year for reinvestment.
  4. Risk concentration. Credit-card debt is uncapped-rate, demand debt. Holding brokerage cash yielding only a few dollars of taxable interest while carrying a 25.74% balance is negative carry.
  5. Exception. If Iggy can identify a specific, high-confidence flip with a quick, documented margin materially above 25.74% and a buyer already lined up, a small inventory purchase could be justified. The default position, however, is debt reduction first.

Actionable sequence:

  1. Apply $9,000 to the highest-APR cards first.
  2. Keep $1,000–$2,000 of emergency liquidity if any platform balances can be withdrawn without business disruption; otherwise pay all $9,000.
  3. Redirect the freed monthly cash flow (former minimum payments + interest savings) to a "restock" sinking fund until the cards are fully paid off.
  4. Once the remaining ~$3,000 balance is eliminated, resume inventory purchases using cash, not cards.

11. Tolerances and Sensitivity Notes

  1. Rounding. All dollar values are rounded to the nearest cent in the detail tables and to the nearest dollar in the summary. A $1–$5 rounding variance is expected if the same computations are reproduced with intermediate rounding.
  2. Allocation methods. Shipping, supplies, credit-card interest, and bank maintenance fees are allocated by reasonable proxies (order count, purchase category, and deposit share). Changing the allocation method can shift net profit between Schedule Cs by a few hundred dollars but will not change the combined Schedule C net profit or the bottom-line federal tax by more than a few dollars.
  3. QBI. If the §481(a) adjustment were excluded from QBI, the QBI deduction would fall by $840.00 and taxable income / regular tax would rise by roughly $168–$185, increasing the balance due by the same amount.
  4. App startup costs. If the software purchases were currently deducted rather than capitalized, Schedule C would need to be created for an app business (which has no revenue), producing a larger loss and a slightly lower bottom-line tax. The ground truth treats the app as pre-active and capitalizes the costs, consistent with the source statement that it is "half-built."
  5. Standard deduction vs. itemizing. Iggy uses the standard deduction; itemizing would not change the result unless state/local taxes and mortgage interest exceed $15,750, which is not supported by the source documents.

12. Common Model Failure Modes

When evaluating other responses or automated computations for Persona 41, watch for these typical errors.

Failure mode Why it is wrong Correct treatment
Ignoring the §481(a) adjustment The Form 3115 positive adjustment of $4,200 is income in 2025 and also self-employment income. Include it in gross income and in the SE/QBI base.
Placing §481(a) adjustment on the wrong form Some models report it only as "other income" and omit it from Schedule SE and QBI. It should flow through Schedule C or Schedule 1 as business income and be included in SE/QBI because it arises from a business accounting-method change.
Expensing the app software immediately The app is pre-revenue and not yet active. Capitalize $2,640 of software plus $232 of related interest as startup costs under §195; no amortization in 2025.
Treating all credit-card interest as deductible Personal credit-card interest is non-deductible under §163(h). Only $2,214.69 is business interest; $465.31 is personal and non-deductible.
Allocating shipping/supplies 50/50 or ignoring them The sneaker business has far more transactions than Etsy. Allocate by order count: Etsy 92/295, sneakers 203/295.
Forgetting the standard deduction 2025 single standard deduction is $15,750, not the 2024 amount. Subtract $15,750 from AGI.
Using 2024 brackets or rates The user explicitly asked for 2025 federal tax. Use 2025 single brackets and 2025 Social Security wage base ($176,100).
Double-counting W-2 Social Security tax Some models subtract withheld SS tax from the SE tax base. Withheld Social Security tax is not a credit against SE tax; it simply reduces how much of the wage base remains for the 12.4% portion.
Incorrect QBI cap QBI deduction cannot exceed 20% of taxable income before the deduction. Tentative QBI = $4,034.01; cap = $13,799.86; allowed = $4,034.01.
Using $9,000 as brokerage cash The brokerage statement shows $3,800; the $9,000 includes platform balances. Use $3,800 for brokerage liquidity and $9,000 only for total available liquid funds when analyzing debt pay-down.
Creating three Schedule Cs The app has no revenue and no active business. Two Schedule Cs: Etsy and combined sneaker reselling. App costs are capitalized.
Mixing up gross vs. net platform receipts Some models use gross platform receipts without subtracting refunds or fees. Net receipts = gross − refunds; then subtract fees, COGS, and allocated operating expenses.

13. Bottom-Line Summary

Metric Value
Combined Schedule C net profit $15,970
§481(a) adjustment $4,200
Self-employment tax $2,850
Deductible half of SE tax $1,425
QBI deduction $4,034
2025 federal taxable income $64,965
Regular federal income tax $9,206
Total 2025 federal tax liability $12,056
Federal tax withheld $6,120
Balance due with return $5,936
Etsy verdict Profitable ($4,824, ~39% margin)
Sneaker verdict Profitable ($11,146, ~33% margin)
App verdict Pre-revenue; $2,872 capitalized
Credit-card recommendation Pay down $9,000 of the $12,000 balance before buying more inventory

Prepared by: Finance Verifier Agent — Persona 41

Date: 2026-07-11

Sources: the private evaluation source set and the 12 constituent source files listed in §1.

Task 03

Bonus Tax & Debt Allocation

Tax planning · Credit decision · Goal-based planning

Two W-2s, a signing bonus and a year-end bonus. How much survives federal tax, and what debt should it kill first?

6 documents

The PromptWhat the simulated consumer asks, in their own voice+

Persona 34 — Kwame Adjei, Task 4: Bonus Tax and Debt Allocation

Persona voice: 25-year-old single investment-banking analyst in Chicago, first real money year, anxious but detail-hungry, tired of being told "pay off debt" without seeing the actual numbers. Wants to know exactly how much of his year-end bonus survives taxes and what it can actually kill.

Assigned category: Tax planning, Credit decision, Goal-based planning


The Prompt

Hey, I'm Kwame Adjei, I'm 25, single, living in Chicago, and I just finished my first full year as an investment banking analyst at Halstead Crane. I started in June 2025 after leaving Northbridge in May, so I have two W-2s and I feel like my tax situation got weird. I did my own 2025 federal return in a free tool and it says I owe about $4,703, but I don't trust it because of the two jobs, the signing bonus, and the year-end bonus. I need you to check that number, tell me the exact corrected balance due, and show me the exact dollar adjustment from my draft. Don't just say it "looks fine" — I want to see the math.

My Halstead compensation for 2025 was $110,000 base, a $25,000 signing bonus, and a $79,000 year-end bonus. The year-end bonus should show up in the 2025 Halstead W-2 pay-period register. I've uploaded both W-2s, my 2026 Halstead paystub, and my bank statements. I'm single, federal only, and I want 2025 numbers throughout.

Here's what I actually need to decide. I've been telling myself I'm going to throw my bonus money at my 9% Grad PLUS loan because that rate is brutal, but I also have this credit card balance that keeps going up even though I put it on autopay, and I'm scared the IRS bill is going to eat way more than I think. So I need to know: after all applicable federal taxes, how much of my $104,000 total bonus is actually left? Not a range — I want the real number and I want to know how you computed it, because I've read conflicting things about marginal versus average tax rates on bonuses and I don't know which one applies here.

Then I need you to solve the debt priority question for me. I have about $165,000 in student loans and a high-APR credit card. Pull the actual balances and rates from the statements I uploaded and list every loan group. I want to know whether I should kill the credit card first or throw the bonus straight at the 9% Grad PLUS, and I want to see the rate math that proves it.

Finally, I need a real allocation plan. I want to keep at least $6,000 combined in checking and savings, I do not want to sell my VTI in the brokerage, I do not want to pause or reduce my 401(k) deferrals, and I do not want the main answer to be "hire a CPA." Show me exactly what dollars go where — the corrected tax bill, the credit card, the Grad PLUS, and what remains. And if there's something bigger I'm missing in these documents, I want you to name it, show me the numbers, and explain why it changes the plan. Cite the actual file and field for every dollar figure so I can follow it myself.


Files the model should expect to reason across

  1. Northbridge 2025 W-2
  2. Halstead Crane 2025 W-2 (includes 2025 pay-period register)
  3. 2026 Halstead paystub
  4. 2026 checking and savings statements
  5. Student loan servicer statement
  6. Credit card statement
  7. Draft 2025 Form 1040 / tax summary

Prompt checklist

  • Consumer live worry: how much of the bonus survives taxes and what it can kill.
  • Concrete source-specific facts in the user's voice.
  • Hidden correction: the credit card is negatively amortizing and is the bigger trap, not the tax bill or the 9% loan.
  • Requires cross-file reasoning across 5+ documents.
  • Method guardrails implied (federal-only, 2025 numbers, cite file/field, preserve constraints).
  • No forbidden escape hatch presented as the main answer.
  • Demands exact dollar adjustments, not ranges.
  • Preserves all persona constraints: $6k cash floor, VTI untouched, 401(k) untouched, no CPA as main answer.
Critical ElementsThe findings a correct response must reach+

Persona 34 — Kwame Adjei, Task 4: Critical Elements

  1. The hidden headline is that the 24.99% credit card is negatively amortizing and must be paid before any student-loan prepayment. The prompt frames the problem around taxes and the 9% Grad PLUS, but the credit-card statement shows a $3,383.83 balance at 24.99% with monthly interest of about $65.59 exceeding the autopay amount of about $63.08, so the balance grows despite payments. The correct response must reframe this as the most urgent trap and prove it with the statement's minimum-payoff warning of 23 years 4 months and $12,744.61 before addressing the tax bill or loans. This maps to Positive #1 (hidden headline), Positive #10–#11 (credit-card mechanics and payoff warning), and Negative #38 (misses the headline or pays loans before the card).
  1. The corrected 2025 federal balance due is $3,494.17, which is $1,208.83 lower than Kwame's draft $4,703.00. The draft tracker is wrong because it kept the $2,500 student-loan interest deduction, missed the excess Social Security credit, and ignored the Additional Medicare Tax. Reconciling both W-2s with the 2025 single standard deduction, the disallowed student-loan deduction, the $55.50 Additional Medicare Tax, and the $1,864.13 excess SS credit produces the corrected bottom line and the exact downward adjustment from the draft. This maps to Positive #2 (corrected balance), Positive #3 (adjustment), Positive #18 (standard deduction), and Negative #37 (wrong standard deduction or non-2025 constants).
  1. The 2025 single standard deduction is $15,750 and the student-loan interest deduction is $0 because MAGI of roughly $204,787 clears the $100,000 phase-out ceiling. Kwame's MAGI sits well above the 2025 single student-loan phase-out range of $85,000 to $100,000, so the full $2,500 draft deduction must be disallowed. The response must use $15,750 and explicitly state the student-loan deduction is $0, not silently keep the draft's $2,500. This maps to Positive #16 (MAGI above ceiling), Positive #17 (SL deduction $0), Positive #18 (standard deduction), and Negative #37 (wrong standard deduction or non-2025 constants).
  1. The two-job situation produces a $1,864.13 excess Social Security credit and $55.50 of Additional Medicare Tax. Combined Social Security wages of $206,166.67 exceed the 2025 wage base of $176,100, so the over-withheld Social Security tax becomes a refundable credit of $1,864.13. Combined Medicare wages of $206,166.67 exceed the $200,000 single threshold, triggering $55.50 of Additional Medicare Tax owed. Both figures must appear in the reconciliation; omitting either breaks the balance due. This maps to Positive #19 and Positive #20.
  1. Halstead 2025 gross compensation is $168,166.67 before the $4,000 401(k) deferral, giving Box 1 of $164,166.67. The new brief is authoritative: seven months of $110k base ($64,166.67) plus a $25,000 signing bonus plus a $79,000 year-end bonus totals $168,166.67. The Halstead W-2 confirms this through the $168,166.67 Box 3/5 totals and the Dec 30 pay-period spike to $83,583.38. With the $4,000 Box 12a Code D 401(k) deferral, Box 1 is $164,166.67, not the gross amount. This maps to Positive #6 (total bonus) and Positive #21 (compensation reconstruction).
  1. The after-tax bonus from the $104,000 total bonus is approximately $73,300 under the strict "all applicable federal taxes" interpretation. The prompt now asks for the amount left after all applicable federal taxes. The strict net-paycheck method subtracts every employee-paid federal tax attributable to the bonus: incremental federal income tax ($24,543.73), net Social Security tax after the excess-credit interaction ($4,583.87), and total Medicare tax including Additional Medicare ($1,563.50), for total federal tax on the bonus of $30,691.10. Bonus left = $104,000.00 − $30,691.10 = $73,308.90, presented as ~$73,300 with a tolerance band of $72,500–$75,500. The response must show this strict method and explicitly reject the 22% supplemental withholding shortcut. This maps to Positive #4 (bonus-after-tax figure) and Positive #5 (method shown).
  1. Debt priority is the credit card at 24.99% first, then Group C Grad PLUS at 9.00%; the other student-loan groups are lower priority. The credit card at $3,383.83 and 24.99% is the highest-cost debt. Among student loans, Group C Grad PLUS at $45,444.80 and 9.00% is the worst, followed by Group D at 7.12%, Group B at 6.54%, and Group A at 4.99%. The response must rank them by APR and direct every extra dollar to the card before any student-loan prepayment. This maps to Positive #7 (card above all loans), Positive #8 (loan-group order), Positive #9 (card balance/APR), and Negative #34 (bonus to Grad PLUS while card still carries a balance).
  1. The bonus can fully eliminate the $45,444.80 Grad PLUS while preserving the $6,000 cash floor. Two paths are defensible:
  • Bonus-only path (Scenario A): after-tax bonus $73,308.90 − corrected tax $3,494.17 − credit card $3,383.83 − cash floor $6,000.00 − Grad PLUS $45,444.80 = $14,986.10 surplus.
  • Recommended cash-first path (Scenario B): pay corrected tax and credit card from existing $20,862.28 liquid cash ($13,984.28 remains above the floor), then deploy the entire after-tax bonus to Grad PLUS: $73,308.90 − $45,444.80 = $27,864.10 surplus.

Either path proves the bonus is large enough to wipe the 9% group without violating the floor. This maps to Positive #12 (bonus pays Grad PLUS), Positive #13 (cash floor preserved), Positive #14 (bonus-only surplus), Positive #15 (cash-first surplus), and Positive #28 (sequenced plan).

  1. The sequenced allocation plan must preserve all hard constraints: $6,000 cash floor, no VTI sale, no 401(k) reduction or pause, and no CPA as the main answer. Kwame's constraints are non-negotiable. The response must not recommend selling the brokerage VTI, pausing or reducing 401(k) contributions, or outsourcing the core tax and allocation math to a CPA. The plan must leave at least $6,000 combined in checking and savings after the allocation. This maps to Positive #28 (sequenced plan), Negative #30 (CPA), Negative #31 (cash floor), Negative #32 (VTI sale), and Negative #33 (401(k) reduction).
  1. Every major number must be tied to a named source document and field. Kwame explicitly asked for file-and-field citations, so the response must name the source when each key figure is introduced: the Northbridge W-2 Box 1 wage of $38,000, the Halstead Crane W-2 Box 1 wage of $164,166.67, the $4,000 Box 12a Code D 401(k) deferral, the combined Box 2 federal withholding of $32,313, the $120 1099-INT interest, the $4,703 draft balance due, and the $3,494.17 corrected balance due should all trace to the tax documents; the credit-card balance of $3,383.83, APR of 24.99%, statement minimum payment of $67.68, and minimum-payoff warning of 23 years 4 months / $12,744.61 should all be tied to the credit-card statement; the Group A through D student-loan balances of $40,610.25, $56,080.83, $45,444.80, and $17,404.38 and their rates of 4.99%, 6.54%, 9.00%, and 7.12% should come from the student-loan statement; and the combined liquid cash of $20,862.28 should come from the checking and savings statements. This maps to Positive #27 (document sourcing).
  1. The analysis must be federal-only and use 2025 constants throughout. The prompt restricts the analysis to federal tax and debt allocation. The response must not bring in Illinois state tax, Chicago local tax, or non-2025 brackets or standard deductions. Required constants are the 2025 single standard deduction of $15,750, the 2025 Social Security wage base of $176,100, the 2025 Additional Medicare Tax single threshold of $200,000, and the 2025 single ordinary-income brackets. Net Investment Income Tax is not included in the project-source federal tax chain because the source documents do not state the 2025 NIIT threshold or rate. This maps to Positive #29 (federal-only scope) and Negative #37 (wrong standard deduction or non-2025 constants).
Golden TrajectoryStep-by-step path to the answer, every figure sourced+
  1. Navigate to W2_Northbridge_2025.pdf and retrieve all wage and withholding boxes: Box 1 wages $38,000.00, Box 2 federal withholding $4,200.00, Box 3 Social Security wages $38,000.00, Box 4 Social Security tax withheld $2,356.00, Box 5 Medicare wages $38,000.00, Box 6 Medicare tax withheld $551.00. 2. Navigate to W2_HalsteadCrane_2025.pdf and retrieve all wage and withholding boxes: Box 1 wages $164,166.67, Box 2 federal withholding $28,113.00, Box 3 / Box 5 wages $168,166.67, Box 4 Social Security tax withheld $10,426.33, Box 6 Medicare tax withheld $2,438.42, Box 12a Code D 401(k) deferral $4,000.00. 3. Calculate the combined W-2 totals by adding the corresponding boxes from Step 1 and Step 2: Box 1 wages $38,000.00 + $164,166.67 = $202,166.67, Box 3/5 wages $38,000.00 + $168,166.67 = $206,166.67, Box 2 federal withholding $4,200.00 + $28,113.00 = $32,313.00, Box 4 Social Security tax $2,356.00 + $10,426.33 = $12,782.33. 4. Navigate to 1099INT_Lakeshore_2025.pdf, go to Box 1, and retrieve taxable interest of $120.00. 5. Calculate AGI by adding the combined Box 1 wages from Step 3 and the interest income from Step 4: $202,166.67 + $120.00 = $202,286.67. 6. Navigate to tax_tracker_DRAFT_2025.xlsx, go to the 1040_Draft sheet, and retrieve: draft balance due $4,703.00, and draft student-loan interest deduction $2,500.00. 7. Calculate MAGI for the student-loan interest deduction by adding AGI from Step 5 and the draft student-loan deduction from Step 6: $202,286.67 + $2,500.00 = $204,786.67. 8. Use the 2025 single student-loan interest phase-out range of $85,000–$100,000. Because MAGI from Step 7 exceeds the $100,000 ceiling, the allowed student-loan interest deduction is $0. 9. Use the 2025 single standard deduction of $15,750. 10. Calculate taxable income by subtracting the allowed student-loan deduction from Step 8 and the standard deduction from Step 9 from AGI in Step 5: $202,286.67 − $0 − $15,750.00 = $186,536.67. 11. Calculate regular federal income tax by applying the 2025 single ordinary brackets to the taxable income from Step 10: 10% bracket $11,925 × 10% = $1,192.50, 12% bracket $36,550 × 12% = $4,386.00, 22% bracket $54,875 × 22% = $12,072.50, 24% bracket $83,186.67 × 24% = $19,964.80, total regular tax = $37,615.80. 12. Calculate the excess Social Security credit. Using the 2025 Social Security wage base of $176,100, the maximum Social Security tax is $176,100 × 6.2% = $10,918.20. Subtract this from the total Social Security tax withheld from Step 3: $12,782.33 − $10,918.20 = $1,864.13. 13. Calculate the Additional Medicare Tax using the combined Medicare wages from Step 3 and the 2025 single threshold of $200,000: ($206,166.67 − $200,000.00) × 0.9% = $55.50. 14. Calculate the corrected federal balance due by adding the regular tax from Step 11 and the Additional Medicare Tax from Step 13, subtracting the excess Social Security credit from Step 12, then subtracting combined federal withholding from Step 3: ($37,615.80 + $55.50 − $1,864.13) − $32,313.00 = $3,494.17. 15. Calculate the adjustment versus the draft by subtracting the corrected balance due from Step 14 from the draft balance due from Step 6: $4,703.00 − $3,494.17 = −$1,208.83. 16. Calculate total bonus: $25,000 signing + $79,000 year-end = $104,000.00. Then calculate the after-tax bonus under the strict "all applicable federal taxes" interpretation by subtracting every employee-paid federal tax attributable to the bonus from the gross bonus: incremental regular federal income tax on the bonus = $37,615.80 − $13,072.07 = $24,543.73; Social Security tax withheld on the bonus = $104,000.00 × 6.2% = $6,448.00; excess SS credit attributable to the bonus = $1,864.13; net Social Security tax on the bonus = $6,448.00 − $1,864.13 = $4,583.87; regular Medicare tax on the bonus = $104,000.00 × 1.45% = $1,508.00; Additional Medicare Tax on the bonus = $55.50; total Medicare tax on the bonus = $1,508.00 + $55.50 = $1,563.50; total federal tax attributable to the bonus = $24,543.73 + $4,583.87 + $1,563.50 = $30,691.10; after-tax bonus = $104,000.00 − $30,691.10 = $73,308.90, presented as ~$73,300 with tolerance $72,500–$75,500. 17. Navigate to credit_card_statement_2026.pdf and retrieve: new balance $3,383.83, APR 24.99%, minimum-payoff warning 23 years 4 months / $12,744.61, average monthly interest ~$65.59 and average monthly autopay ~$63.08 from year-to-date totals. Conclude that the credit card is negatively amortizing because monthly interest exceeds autopay; this is the hidden headline. 18. Navigate to student_loan_statement_2026.pdf and retrieve the loan groups: Group A $40,610.25 @ 4.99%, Group B $56,080.83 @ 6.54%, Group C (Grad PLUS) $45,444.80 @ 9.00%, Group D $17,404.38 @ 7.12%. 19. Establish debt priority by APR: credit card 24.99% first, then Group C Grad PLUS 9.00%, then Group D 7.12%, then Group B 6.54%, then Group A 4.99%. 20. Navigate to bank_checking_2026.pdf and retrieve the ending checking balance of $10,010.92; navigate to savings_statement_Lakeshore_2026.pdf and retrieve the ending savings balance of $10,851.36. 21. Calculate combined liquid cash by adding the balances from Step 20: $10,010.92 + $10,851.36 = $20,862.28. 22. Apply the user's stated constraints: preserve at least $6,000 combined liquid cash, do not sell VTI, do not pause or reduce 401(k) deferrals, and do not make "hire a CPA" the main answer. 23. Build the bonus-only allocation path. Start with the after-tax bonus from Step 16 and pay the immediate obligations from it: pay corrected 2025 federal tax from Step 14 −$3,494.17, pay credit-card balance in full from Step 17 −$3,383.83, preserve the $6,000 cash floor from Step 22 −$6,000.00, pay off Grad PLUS Group C from Step 18 −$45,444.80; remaining bonus cash = $73,308.90 − $3,494.17 − $3,383.83 − $6,000.00 − $45,444.80 = $14,986.10 surplus. 24. Build the recommended cash-first allocation path. Pay the corrected tax and credit-card balance from the existing liquid cash in Step 21: $20,862.28 − $3,494.17 − $3,383.83 = $13,984.28, which is above the $6,000 floor from Step 22. Then deploy the entire after-tax bonus from Step 16 against Grad PLUS Group C from Step 18: $73,308.90 − $45,444.80 = $27,864.10 surplus. 25. State the final verified answer: corrected 2025 federal balance due $3,494.17, which is $1,208.83 lower than the draft $4,703.00; after-tax bonus available ~$73,300; the bonus is enough to fully eliminate the $45,444.80 Grad PLUS Group C while preserving the $6,000 cash floor, leaving roughly $27,900 surplus under the recommended cash-first plan or roughly $14,900 under the bonus-only plan; the hidden headline is that the 24.99% credit card is the most urgent trap because it is negatively amortizing on autopay minimums.
Scoring RubricWeighted positive and negative criteria+

Persona 34 — Kwame Adjei, Task 4: Rubric

Total positive points: 795

Total negative points: −430 (applied only when triggered)

Criteria count: 35

This rubric was tightened after sequential verification of three model responses and a program auto-rubric review. The final design resolves item-atomicity, MECE, positive-phrasing, objectivity, and prompt-coverage issues while keeping weight on the criteria that actually separate pass from fail on this task: the hidden headline, the exact corrected tax balance, credit-card/Grad-PLUS priority, and proof that the bonus can fully eliminate the $45,444.80 Grad PLUS loan.


Positive rubric items

# Item Points Pass criterion Rationale
1 The response states the corrected 2025 federal balance due is exactly $3,494.17, a $1,208.83 reduction from the draft $4,703.00. +100 Response gives the exact corrected federal balance due and the exact dollar adjustment from the draft. The consumer explicitly asks for both the corrected balance and the adjustment; the two figures are arithmetically dependent on the same tax-chain result.
2 The response states the after-tax value of the $104,000 total bonus is about $73,300, within the $72,500–$75,500 tolerance band. +30 Response places the after-tax bonus inside the $72,500–$75,500 band and ties it to the $104,000 gross bonus. The consumer wants the real after-tax bonus number; the ground-truth strict all-federal-taxes figure is about $73,300.
3 The response applies marginal federal income-tax brackets to compute the bonus after-tax figure and explicitly rejects the 22% supplemental withholding shortcut. +45 Response shows the bonus-after-tax method using incremental regular tax and states the 22% withholding rate is not the final tax rate. Without this guardrail, models quote withholding as if it were tax and produce the wrong net bonus.
4 The response states the total bonus is $104,000, comprising the $25,000 signing bonus and the $79,000 year-end bonus verified in 2025 Halstead compensation. +30 Response reconciles the $25,000 signing bonus and the $79,000 year-end bonus to $104,000 total and confirms the year-end bonus in Halstead pay records. The brief’s bonus chain is $25,000 signing + $79,000 year-end = $104,000; the consumer specifically wants the year-end bonus verified.
5 The response ranks the 24.99% credit card above every student loan for priority payment. +50 Response explicitly states the credit card must be paid before any student loan. The consumer asks whether to kill the card first or throw the bonus at the 9% Grad PLUS; the rate math makes the credit card the clear first priority.
6 The response identifies the 9.00% Grad PLUS as the highest-interest student-loan bucket. +30 Response states the 9.00% Grad PLUS is the worst student loan by rate. The consumer wants the rate math ranking loans; the 9.00% Grad PLUS is the highest-rate student-loan group.
7 The response ranks student-loan prepayment priority as 9.00%, then 7.12%, then 6.54%, then 4.99%. +15 Response orders the four loan groups by rate: $45,444.80 at 9.00% first, $17,404.38 at 7.12% second, $56,080.83 at 6.54% third, $40,610.25 at 4.99% last. The consumer wants the rate math ranking loans; this is the correct descending-rate order from the loan statement.
8 The response states the credit card balance is $3,383.83. +10 Response gives the exact $3,383.83 balance from the credit-card statement. The consumer asks for actual balances; the credit card is the highest-cost debt.
9 The response states the credit card APR is 24.99%. +10 Response gives the exact 24.99% APR from the credit-card statement. The APR drives the negative-amortization trap.
10 The response shows monthly credit-card interest of about $65.59 exceeds the autopay amount of about $63.08, so the balance grows. +20 Response computes interest of about $65.59 versus autopay of about $63.08 and concludes the card is negatively amortizing. The consumer noticed the balance rising despite autopay; the proof must come from the statement numbers.
11 The response quotes the credit-card minimum-payoff warning of 23 years 4 months and $12,744.61 total. +15 Response quotes the statement warning of 23 years 4 months and $12,744.61 minimum-payoff cost. The statement warning proves the long-term cost of the negative-amortization trap.
12 The response recommends applying after-tax bonus funds to fully pay off the $45,444.80 Grad PLUS loan. +70 Response directs the after-tax bonus to eliminate the $45,444.80 Grad PLUS group and demonstrates the resulting surplus. The consumer’s stated plan is to throw the bonus at the 9% Grad PLUS; the response must prove it can be fully paid without breaking the cash floor.
13 The response shows the allocation leaves at least $6,000 combined in checking and savings. +40 Response demonstrates the plan preserves the $6,000 combined cash floor. The consumer sets a hard constraint of keeping at least $6,000 combined in checking and savings.
14 The response states the bonus-only surplus after tax, credit card, Grad PLUS, and cash floor is about $14,986.10. +15 Response gives the bonus-only residual of about $14,986.10. The consumer asks what remains after the allocation; the bonus-only path leaves about $14,986.10.
15 The response states MAGI is roughly $204,787, above the $100,000 single ceiling, so the 2025 student-loan interest deduction is $0. +70 Response gives MAGI of roughly $204,787, notes it clears the $100,000 single phase-out ceiling, and sets the deduction to $0. The rule and the answer are the same correction; a response must show both the MAGI figure and the resulting $0 deduction.
16 The response includes a refundable excess Social Security credit of $1,864.13. +45 Response includes the $1,864.13 refundable excess Social Security credit. Combined Social Security wages exceed the $176,100 wage base, producing this credit.
17 The response includes $55.50 of Additional Medicare Tax owed. +25 Response includes $55.50 of Additional Medicare Tax. Combined Medicare wages exceed the $200,000 single threshold, triggering $55.50.
18 The response states the Halstead Crane W-2 Box 1 wage is $164,166.67 after the $4,000 Box 12a Code D 401(k) elective deferral. +35 Response gives Halstead Box 1 of $164,166.67 and identifies the $4,000 Box 12a Code D 401(k) deferral. Core input to the tax reconciliation; gross compensation minus the $4,000 401(k) deferral yields the Box 1 taxable wage.
19 The response states the Northbridge W-2 Box 1 wage is $38,000. +10 Response gives the Northbridge Box 1 wage of $38,000. Required input reconciling both jobs in the tax chain.
20 The response states combined federal income tax withholding from Box 2 of both W-2s is $32,313.00. +10 Response gives combined Box 2 federal withholding of $32,313.00. Required figure for the balance-due calculation.
21 The response includes $120.00 of 1099-INT interest income. +10 Response includes $120.00 of taxable interest income in AGI. The 1099-INT is part of the income picture.
22 The response states combined liquid cash across checking and savings is $20,862.28. +10 Response gives combined checking and savings cash of $20,862.28. Supports the alternative allocation path and the $6,000 floor check.
23 The response states loan Group A balance is $40,610.25 at 4.99% APR. +5 Response gives the exact Group A balance and rate. The consumer asks to list every loan group with actual balances and rates.
24 The response states loan Group B balance is $56,080.83 at 6.54% APR. +5 Response gives the exact Group B balance and rate. Same as above.
25 The response states loan Group C balance is $45,444.80 at 9.00% APR. +5 Response gives the exact Group C balance and rate. Same as above; this is the Grad PLUS group.
26 The response states loan Group D balance is $17,404.38 at 7.12% APR. +5 Response gives the exact Group D balance and rate. Same as above.
27 The response cites a source document and field for each of the following dollar figures: W-2 Box 1 wages, 401(k) deferral, combined Box 2 withholding, 1099-INT interest, draft and corrected balance due, credit-card balance/APR/payoff warning, each loan-group balance/rate, checking balance, savings balance, and the after-tax bonus. +40 Response names the source document and field for each enumerated figure. The consumer explicitly asks for the actual file and field behind every dollar figure; the list removes grader subjectivity.
28 The response presents a sequenced allocation plan covering corrected tax, credit card, cash floor, Grad PLUS, and remaining surplus. +20 Response shows a phased plan in order. The consumer asks for a real allocation plan showing dollars going where.
29 The response limits all tax calculations and figures to the federal level; no Illinois state or Chicago local tax is included. +20 Response excludes state and local tax from the analysis. The consumer restricts the analysis to federal only.

Negative rubric items

# Item Penalty Trigger Rationale
30 The response recommends pausing or reducing 401(k) contributions. −70 Response suggests pausing or reducing 401(k) deferrals. The consumer explicitly forbids touching 401(k) contributions.
31 The response gives a range, an "it depends" answer, or otherwise refuses to commit to specific dollar figures. −40 Response refuses to commit to a number despite sufficient documents. The consumer demands real numbers, not ranges.
32 The response uses a wrong standard deduction, wrong tax-year brackets, or any non-2025 federal tax constant. −100 Response uses $15,000 standard deduction, 2024 brackets, wrong SS wage base, wrong Additional Medicare threshold, or similar. The 2025 single standard deduction ($15,750), wage base ($176,100), and Medicare threshold ($200,000) are project constants; getting them wrong corrupts the whole tax chain.
33 The response presents the Grad PLUS or another student loan as the top finding without explicitly flagging the negatively amortizing 24.99% credit card as the single highest-priority debt action item. −100 Response names a different "bigger thing missing" or treats a student loan as the headline issue. The consumer explicitly asked what bigger issue he is missing; the credit-card trap is the headline finding.
34 The response recommends paying any student-loan prepayment before the $3,383.83 credit card balance is fully paid off. −50 Response recommends paying a student loan while the credit card still has a balance. Violates the rate-priority reasoning the consumer asks for.
35 The response recommends selling or liquidating VTI holdings in the brokerage account to fund any part of the allocation plan. −70 Response suggests selling brokerage VTI holdings. The consumer explicitly forbids selling VTI.

Score interpretation

Score Verdict
700–795 Excellent — full hidden headline, correct math, clean sourcing, all constraints respected.
600–699 Good — headline and main math correct, minor sourcing or formatting gaps.
500–599 Marginal — main tax or bonus number roughly right but misses headline or one major constraint.
400–499 Poor — surface details right but fails headline, debt priority, or a hard constraint.
0–399 Unacceptable — major errors, rubber-stamps draft, or violates multiple constraints.

Pass/fail threshold for this task: A response must score at least 398 / 795 (50% of the total positive points) to pass.

Model ScoringHow each model response scored, and why+

Persona 34 — Kwame Adjei, Task 4: Response Quality Summary

Scored against the final 35-item program-resolved rubric (persona_34_task4_rubric.md):

Response Score Verdict band Pass/fail (≥398/795)
Response 1 290 / 795 Unacceptable FAIL
Response 2 390 / 795 Poor FAIL (by 8 points)
Response 3 610 / 795 Good PASS

The final rubric resolves every program auto-rubric feedback item: item-atomicity splits, MECE merges, positive-phrasing rewrites, objective verifiability, source-list enumeration, and VTI coverage. It keeps weight on the criteria that separate the responses: the hidden headline, the exact corrected tax balance, credit-card/Grad-PLUS priority, and proof that the bonus fully eliminates the $45,444.80 Grad PLUS loan. The pass threshold is set at 50% of the total positive points (398 / 795).


Response 1 — FAIL (290 / 795)

Response 1 is surface-correct on many inputs but fails the rubric's highest-weighted deliverables.

Critical failures

  • Inverts the 2025 single standard deduction. The response claims $15,750 is wrong and applies $15,000 instead. The project constant is $15,750. This single error inflates taxable income by $750 and regular tax by $180, producing a corrected balance due of $3,674.17 instead of $3,494.17 and an adjustment of −$1,028.83 instead of −$1,208.83. It loses Positive #1 (+100) and triggers Negative #32 (−100).
  • Misidentifies the hidden headline. The response labels the excess Social Security credit as the "bigger thing missing" instead of the negatively amortizing 24.99% credit card. It triggers Negative #33 (−100).
  • No credit-card negative-amortization proof. The response says "kill the credit card first" but never shows that monthly interest exceeds autopay or quotes the statement's 23 years 4 months / $12,744.61 payoff warning. It loses Positive #10 (+20) and Positive #11 (+15).
  • Does not prove the bonus can eliminate the Grad PLUS. The allocation plan sends only $7,804.28 to Group C, leaving $37,640.52 unpaid, and never reports the required bonus-only (~$14,986) surplus. It loses Positive #12 (+70) and Positive #14 (+15).
  • Missing MAGI / $0 deduction. The response uses AGI ($202,286.67) to disallow the student-loan deduction but never states MAGI (~$204,787). It loses Positive #15 (+70).

Strengths

  • W-2 Box 1 values, combined withholding ($32,313), 1099-INT ($120), excess SS credit ($1,864.13), and Additional Medicare Tax ($55.50) are correct.
  • After-tax bonus $73,293.90 lands inside the strict-method tolerance band.
  • Lists all four loan groups with balances/rates and preserves all hard constraints.

Bottom line: Response 1 fails because it corrupts the tax chain with a wrong constant, misses the central hidden headline, and cannot prove the bonus eliminates Grad PLUS. A careful reviewer would refuse to ship this.


Response 2 — FAIL (390 / 795)

Response 2 gets the hidden headline right and sources most figures, but it falls 8 points below the 398 threshold because the standard-deduction error and the Grad-PLUS payoff failure carry enough combined weight.

Critical failures

  • Uses the wrong 2025 single standard deduction ($15,000 instead of $15,750). This inflates the corrected balance due to $3,674.17 instead of $3,494.17 and the adjustment to −$1,028.83 instead of −$1,208.83. It loses Positive #1 (+100) and triggers Negative #32 (−100).
  • Does not prove the bonus can eliminate the Grad PLUS. The allocation plan deploys only $7,803.45 of existing cash to Group C, leaving $37,641.35 unpaid. It never directs the ~$73,300 after-tax bonus to the loan or reports the required surplus figure. It loses Positive #12 (+70) and Positive #14 (+15).
  • Missing supporting items. No monthly negative-amortization proof ($65.59 vs. $63.08), no payoff-warning quote, no MAGI ~$204,787 reasoning, no explicit $0 student-loan interest deduction, and no explicit rate-descending ranking of the four loan groups.

Strengths

  • Correctly identifies the negatively amortizing credit card as the hidden headline and prioritizes it over the Grad PLUS.
  • After-tax bonus $73,293.90 is within tolerance.
  • Respects all hard constraints (cash floor, no VTI, no 401(k)).
  • Strong document sourcing for most figures.

Bottom line: Response 2 is the near-miss. It earns +490 / 795 on the positive rubric by getting the headline, bonus gross, and most tax pieces right, but two high-weight failures drop it below the 398 pass threshold. First, it uses $15,000 instead of the 2025 single $15,750 standard deduction, which loses Positive #1 (+100) and triggers Negative #32 (−100). Second, it never proves the ~$73,300 after-tax bonus can fully eliminate the $45,444.80 Grad PLUS — only $7,803.45 of existing cash is sent to that loan — so it loses Positive #12 (+70) and Positive #14 (+15). The arithmetic is 490 − 100 = 390, which is 8 points below 398. The rubric reallocation intentionally moved weight onto these exact separating criteria, turning a response that would have barely passed into a clear but close fail. It needs rework on the tax constant and the allocation proof before shipping.


Response 3 — PASS (610 / 795)

Response 3 is the strongest of the three. It gets the headline, the after-tax bonus, all 2025 federal constants, and the debt priority right, but it still loses heavily on the exact tax balance and the Grad PLUS payoff proof.

Critical failures

  • Misses the exact corrected balance due. The response reports "≈ $3,500" instead of the exact $3,494.17 and "−$1,203" instead of −$1,208.83. The root cause is an unsupported $5 NIIT that the project source documents do not authorize. It loses Positive #1 (+100).
  • Does not prove the bonus can fully eliminate the Grad PLUS. The allocation table sends only $7,978 of existing cash to Group C and never deploys the ~$73,300 after-tax bonus to the $45,444.80 loan. It also omits the required bonus-only surplus figure. It loses Positive #12 (+70) and Positive #14 (+15).

Strengths

  • Correctly names the 24.99% credit card as the hidden headline and provides negative-amortization proof plus the exact payoff-warning quote.
  • After-tax bonus $73,309 matches the strict all-federal-taxes ground truth.
  • Uses the correct $15,750 standard deduction and all 2025 federal constants.
  • Ranks the loan groups correctly, computes MAGI ~$204,787, disallows the student-loan interest deduction, and preserves every hard constraint.

Bottom line: Response 3 passes with a comfortable margin but is not excellent. The exact tax balance and the Grad PLUS payoff proof are the two remaining gaps that keep it out of the Excellent band. It clears the 398-point pass threshold by 212 points.


Disclaimer: This is a the evaluation program model-evaluation reference sheet derived from the persona brief, project source documents, and the supplied rubric. It is not tax, legal, or financial advice.

Ground TruthVerified reference calculations+

Persona 34 — Kwame Adjei: Task 4 Ground Truth

Verifier role: Finance Verifier Agent, the evaluation program

Authoritative source: persona_34_task4_assigned_brief.md

Verified against: Downloaded persona files in Persona files/ (2025-07-13). Full report in persona_34_task4_verification_report.md.

Cross-checks only (legacy): persona_34_ground_truth_verification.md and persona_34_attempt2_critical_elements.md from Task 2

Regeneration rule: The new brief overrides Task 2 wherever they differ; Task 2 values were used only where the brief was silent before the persona files were downloaded.


1. Assumptions and cross-check notes

Every number below is either taken directly from the Task 4 brief, derived from the brief, or pulled from Task 2 as a cross-check and labeled as such. No external IRS research or web searches were used.

Assumption Value / treatment Source / rationale
Persona 25, single, Chicago, two 2025 W-2s Task 4 brief
Halstead 2025 base $110,000 annual, worked June–December Task 4 brief
Halstead bonuses $25,000 signing + $79,000 year-end = $104,000 total bonus Task 4 brief; the 2025 Halstead W-2 pay-period register confirms this: regular semimonthly gross is $4,583.33, but the 2025-12-30 pay period is $83,583.38, a one-time spike of ~$79,000. The W-2 Box 3/5 totals $168,166.67, which also equals 7 months base ($64,166.67) + $25k signing + $79k year-end.
Halstead 401(k) deferral $4,000 pre-tax Halstead W-2 Box 12a Code D = $4,000.00 (verified against W2_HalsteadCrane_2025.pdf). The 2026 paystub is a different calendar year and should not override the 2025 W-2.
Northbridge wages Box 1 / Box 3 = $38,000; Box 4 = $2,356; Box 2 = $4,200 Northbridge 2025 W-2 (W2_Northbridge_2025.pdf); brief states only "entry-level finance role Jan–May 2025" and gives no dollar amounts.
Halstead federal withholding Box 2 = $28,113 Halstead 2025 W-2 Box 2 (verified against W2_HalsteadCrane_2025.pdf).
1099-INT interest $120 1099INT_Lakeshore_2025.pdf, Box 1 = $120.00.
Credit-card balance, APR, limit, minimum payment, payoff warning $3,383.83; 24.99%; $8,000; $67.68; 23 years 4 months / $12,744.61 credit_card_statement_2026.pdf; brief says only "high-APR credit card."
Student-loan groups and rates A $40,610.25 @ 4.99%; B $56,080.83 @ 6.54%; C $45,444.80 @ 9.00%; D $17,404.38 @ 7.12% student_loan_statement_2026.pdf; brief says only "~$165,000 student loans (one tranche at 9%)."
Combined liquid cash + VTI Checking + savings = $20,862.28; VTI = $1,567.74 bank_checking_2026.pdf ending balance $10,010.92 + savings_statement_Lakeshore_2026.pdf ending balance $10,851.36; brokerage_statement_Fidelity_2026.pdf VTI $1,567.74.
Cash floor Keep ≥ $6,000 liquid User / Task 2 constraint; brief is silent.
401(k), VTI, CPA Do not reduce 401(k), do not sell VTI, do not recommend a CPA User / Task 2 constraint.
Tax constants 2025 single standard deduction $15,750; SS wage base $176,100; Additional Medicare Tax single threshold $200,000; 2025 single ordinary-income brackets Given 2025 IRS constants for this task.

Reconciliation with Task 2: The new brief explicitly states a $79,000 year-end bonus. The actual 2025 Halstead W-2 provided by the platform confirms the total: Box 3/5 = $168,166.67 and the pay-period register shows the 2025-12-30 pay period at $83,583.38 versus a regular $4,583.33, a one-time spike of ~$79,000. Therefore the prompt's $79k figure is not just a user claim; it is supported by the W-2 register and year-end totals.


2. Compensation reconstruction

Halstead 2025 earnings (brief + W-2 register)
Item Formula / source Amount
Base pay, 7 months $110,000 × 7/12 $64,166.67
Signing bonus Brief + W-2 register $25,000.00
Year-end bonus Brief + W-2 register (Dec 30 pay spike) $79,000.00
Gross Halstead compensation before 401(k) Sum $168,166.67
401(k) deferral Halstead W-2 Box 12a Code D −$4,000.00
W-2 Box 1 Gross − deferral; matches W-2 $164,166.67
W-2 Box 3 / Box 5 Gross; matches W-2 $168,166.67
W-2 Box 4 (Social Security tax withheld) $168,166.67 × 6.2%; matches W-2 $10,426.33
W-2 Box 6 (Medicare tax withheld) $168,166.67 × 1.45%; matches W-2 $2,438.42
W-2 Box 2 (federal withholding) Matches W-2 $28,113.00

W-2 register evidence for the $79k bonus:

  • Regular semimonthly gross (Jul–Nov): $4,583.33
  • 2025-12-30 gross: $83,583.38
  • One-time spike: $83,583.38 − $4,583.33 = $79,000.05 ≈ $79,000 year-end bonus
  • Year total gross: $168,166.67 = $64,166.67 base + $25,000 signing + $79,000 year-end

Why Box 1 differs from gross: A pre-tax 401(k) deferral reduces taxable wages in Box 1 but does not reduce Social Security or Medicare wages in Box 3/5.


3. 2025 federal tax chain

3.1 Wage reconstruction
Source Box 1 Box 3 / 5 Box 4 Box 6 Box 2
Northbridge 2025 W-2 $38,000.00 $38,000.00 $2,356.00 $551.00 $4,200.00
Halstead 2025 W-2 $164,166.67 $168,166.67 $10,426.33 $2,438.42 $28,113.00
Combined $202,166.67 $206,166.67 $12,782.33 $2,989.42 $32,313.00
3.2 AGI and deduction test
Item Formula Amount
Combined Box 1 wages Above $202,166.67
1099-INT interest Cross-check +$120.00
AGI $202,166.67 + $120.00 $202,286.67
Draft student-loan interest deduction Draft claim (cross-check) $2,500.00
MAGI for SL deduction AGI + $2,500 $204,786.67
2025 single phase-out range Given constant $85,000 – $100,000
Allowed student-loan deduction MAGI ≥ $100,000 ceiling $0.00
2025 single standard deduction Given constant −$15,750.00
Taxable income $202,286.67 − $15,750.00 − $0 $186,536.67
3.3 Regular federal income tax (2025 single brackets)
Bracket Amount in bracket Rate Tax
$0 – $11,925 $11,925.00 10% $1,192.50
$11,926 – $48,475 $36,550.00 12% $4,386.00
$48,476 – $103,350 $54,875.00 22% $12,072.50
$103,351 – $197,300 $83,186.67 24% $19,964.80
Total regular tax $37,615.80

Formula check: $1,192.50 + $4,386.00 + $12,072.50 + $19,964.80 = $37,615.80.

3.4 Excess Social Security tax credit
Item Formula Amount
Total SS tax withheld (Box 4) $2,356.00 + $10,426.33 $12,782.33
2025 SS wage base Given constant $176,100.00
Maximum SS tax $176,100.00 × 6.2% $10,918.20
Excess SS credit $12,782.33 − $10,918.20 $1,864.13
Overage formula ($206,166.67 − $176,100.00) × 6.2% $1,864.13
3.5 Additional Medicare Tax
Item Formula Amount
Combined Medicare wages (Box 5) Same as combined Box 3 $206,166.67
Single threshold Given constant $200,000.00
Overage $206,166.67 − $200,000.00 $6,166.67
Additional Medicare Tax $6,166.67 × 0.9% $55.50
3.6 Corrected federal balance due
Item Amount
Regular income tax $37,615.80
Additional Medicare Tax $55.50
Excess Social Security credit −$1,864.13
Total corrected federal tax liability $35,807.17
Combined federal withholding (Box 2) −$32,313.00
Corrected federal balance due $3,494.17
Draft balance due (Task 2 cross-check) $4,703.00
Exact adjustment vs. draft −$1,208.83

The draft tracker is $1,208.83 too high because it kept the $2,500 student-loan deduction, missed the excess SS credit, and missed the Additional Medicare Tax. Net investment income tax (NIIT) is not included in this corrected balance because the project source documents do not state the 2025 NIIT threshold or rate, and the $120 of 1099-INT interest is the only investment income item. If an external NIIT constant were imported, the balance would rise by $4.56 to $3,498.73; under project-source-only authority it remains $3,494.17.


4. Bonus-after-tax analysis

Total 2025 Halstead bonus = $25,000 signing + $79,000 year-end = $104,000.

Method 1 — Bracket-only marginal method

Compute only the incremental regular federal income tax attributable to the bonus by stepping through the 2025 single brackets. Non-bonus taxable income is $82,536.67; adding the $104,000 bonus pushes taxable income to $186,536.67.

Portion of bonus Amount Rate Tax
Fills remainder of 22% bracket ($103,350 − $82,536.67) $20,813.33 22% $4,578.93
Falls in 24% bracket ($186,536.67 − $103,350) $83,186.67 24% $19,964.80
Total incremental regular tax on bonus $104,000.00 $24,543.73

Bonus left after regular federal income tax:

$104,000.00 − $24,543.73 = $79,456.27

This method is bracket-only because it ignores FICA interactions (Social Security and Medicare).

Method 2 — Residual / full-marginal method (1040-liability change)

Compute total federal tax liability with the bonus minus total federal tax liability without the bonus.

With-bonus item Amount
Regular income tax $37,615.80
Additional Medicare Tax $55.50
Excess Social Security credit −$1,864.13
Total federal tax liability with bonus $35,807.17
Without-bonus item Formula Amount
Non-bonus Box 1 $202,166.67 − $104,000.00 $98,166.67
AGI without bonus $98,166.67 + $120.00 $98,286.67
Taxable income without bonus $98,286.67 − $15,750.00 $82,536.67
Regular tax without bonus 10% + 12% + 22% brackets $13,072.07
Excess SS credit without bonus SS wages $102,166.67 < base $0.00
Additional Medicare Tax without bonus Medicare wages < $200,000 $0.00
Total federal tax liability without bonus $13,072.07

Tax attributable to the bonus = $35,807.17 − $13,072.07 = $22,735.10

Bonus left after federal income tax and 1040-level payroll effects:

$104,000.00 − $22,735.10 = $81,264.90

Method 3 — Strict “all applicable federal taxes” / net-paycheck method

This method treats the bonus as ordinary wages and subtracts every employee-paid federal tax that applies to it: federal income tax, Social Security tax (net of the excess credit), and Medicare tax (regular + Additional Medicare).

Tax on bonus Formula Amount
Incremental regular federal income tax From Method 1 $24,543.73
Social Security tax withheld on bonus $104,000.00 × 6.2% $6,448.00
Excess SS credit attributable to bonus Already counted in 1040 liability −$1,864.13
Net Social Security tax on bonus $4,583.87
Regular Medicare tax withheld on bonus $104,000.00 × 1.45% $1,508.00
Additional Medicare Tax on bonus Wages over $200,000 attributable to bonus $55.50
Total Medicare tax on bonus $1,563.50
Total federal tax attributable to bonus $24,543.73 + $4,583.87 + $1,563.50 $30,691.10

Bonus left after all applicable federal taxes:

$104,000.00 − $30,691.10 = $73,308.90

Reconciling the three methods
Method Bonus left What it includes
Bracket-only marginal $79,456.27 Federal income tax only
Residual / 1040-liability $81,264.90 Federal income tax + Additional Medicare Tax + excess SS credit interaction
Strict all-federal-taxes $73,308.90 Federal income tax + Medicare (regular + Additional) + net Social Security tax

The residual method is higher than bracket-only because the excess SS credit ($1,864.13) outweighs the Additional Medicare Tax ($55.50). The strict method is lower because it also subtracts the regular Medicare withholding ($1,508.00) and the net Social Security tax ($4,583.87) that the residual method captures only as a 1040-level interaction.

Recommended bonus-after-tax number

The prompt now asks for the amount left after all applicable federal taxes. Under that plain reading, the operative figure is the strict all-federal-taxes result.

  • Recommended: ~$73,300 of the $104,000 bonus remains after all applicable federal taxes.
  • Tolerance band: $72,500 – $75,500 (covers the strict $73,309 estimate and small swings from rounding or alternative payroll assumptions).

Use the strict net-paycheck figure because it matches the prompt language and captures every federal tax that reduces the bonus cash Kwame actually receives.


5. Debt stack and cost-of-debt

Student-loan groups
Group Balance Rate Annual interest
A — Subsidized Stafford $40,610.25 4.99% ~$2,026.45
B — Unsubsidized Stafford $56,080.83 6.54% ~$3,667.69
C — Grad PLUS $45,444.80 9.00% ~$4,090.03
D — Private refinance $17,404.38 7.12% ~$1,239.19
Total student loans $159,540.26 Weighted avg ~6.91% ~$11,023.36

Weighted-average formula:

$$

\frac{(40{,}610.25 \times 0.0499) + (56{,}080.83 \times 0.0654) + (45{,}444.80 \times 0.0900) + (17{,}404.38 \times 0.0712)}{159{,}540.26}

= 0.0691 \approx 6.91\%

$$

Credit card
Item Value
Balance $3,383.83
APR 24.99%
Credit limit $8,000.00
Utilization 42.30%
Statement minimum payment $67.68
Minimum-payoff timeline / total cost 23 years, 4 months / $12,744.61
Debt priority by rate
  1. Credit card at 24.99% — highest-cost debt overall.
  2. Group C Grad PLUS at 9.00% — highest-cost student-loan group.
  3. Group D private refinance at 7.12%.
  4. Group B unsubsidized Stafford at 6.54%.
  5. Group A subsidized Stafford at 4.99%.

6. Liquidity and constraints

Item Amount Source
Combined checking + savings $20,862.28 Task 2 cross-check
VTI brokerage holding $1,567.74 Task 2 cross-check
Hard cash floor ≥ $6,000.00 User constraint
401(k) deferrals Untouched User constraint
CPA recommendation Not allowed User constraint
VTI sale Not allowed User constraint

7. Allocation recommendation

Kwame’s stated goal: throw bonus money at the 9% Grad PLUS. The conservative sequence is:

Scenario A — Bonus-only path

Use only the after-tax bonus for taxes, card, floor, and Grad PLUS.

Step Amount Remaining bonus cash
After-tax bonus $73,308.90 $73,308.90
Pay corrected 2025 federal tax −$3,494.17 $69,814.73
Pay credit-card balance in full −$3,383.83 $66,430.90
Preserve $6,000 combined cash floor −$6,000.00 $60,430.90
Allocate to Grad PLUS Group C −$45,444.80 $14,986.10 leftover

Conclusion: The bonus alone can fully wipe the 9% Grad PLUS and leaves roughly $14,900–$15,000 after also covering the corrected tax bill, the 24.99% credit card, and a fresh $6,000 cash floor.

Scenario B — Recommended path (pay immediate obligations from existing cash)

Pay the corrected tax and credit-card balance from the existing $20,862.28 liquid cash, which preserves the $6,000 floor without touching the bonus.

Step Amount Remaining liquid cash
Start $20,862.28 $20,862.28
Pay corrected 2025 federal tax −$3,494.17 $17,368.11
Pay credit-card balance in full −$3,383.83 $13,984.28

$13,984.28 is well above the $6,000 floor.

Then deploy the entire after-tax bonus to the Grad PLUS:

Bonus deployment Amount
After-tax bonus $73,308.90
Allocate to Grad PLUS Group C −$45,444.80
Surplus after killing Grad PLUS $27,864.10

Recommendation: Use Scenario B. It preserves the cash buffer and puts the maximum possible bonus dollars against the highest-rate student-loan group. The surplus ($27,800+) can then go to the next-highest-rate debt (Group D at 7.12%) or to a down-payment fund, depending on Kwame’s next priority.


8. Hidden headline finding

Even after correcting his 2025 federal tax bill and paying off the 24.99% credit card, Kwame has roughly $73,300 of after-tax bonus cash left — enough to eliminate the entire $45,400 Grad PLUS group in one move while keeping his $6,000 cash floor intact, yet the real financial trap is the negatively amortizing credit card, not the tax bill.


9. Likely model failure modes

A model can get many surface facts right and still fail on the scoring rubric. Watch for these:

  1. Reconstructs Halstead Box 1 as simply base + signing + year-end ($168,166.67) and ignores the $4,000 401(k) deferral, overstating taxable wages by $4,000 and shifting the corrected tax balance by roughly $960.
  2. Uses a flat 24% (or 32%) rate on the entire $104,000 bonus instead of stepping through the 22% and 24% brackets, producing a bracket-only number that is off by several hundred dollars.
  3. Uses the $81,300 residual / 1040-liability method instead of the strict all-applicable-federal-taxes method. The residual method captures the Additional Medicare Tax and excess SS credit interaction but ignores the regular Medicare withholding and net Social Security tax on the bonus, so it overstates bonus cash by about $7,956.
  4. Confuses supplemental 22% withholding with actual tax and reports the bonus "left" as gross minus withholding, which is not after-tax cash.
  5. Forgets the $25,000 signing bonus and analyzes only the $79,000 year-end bonus, understating total bonus cash by $25,000.
  6. Forgets the Additional Medicare Tax on combined Medicare wages above $200,000, producing a balance due ~$55.50 too low.
  7. Misses the excess Social Security credit from two employers, producing a balance due of ~$5,358 and a much smaller bonus-after-tax number.
  8. Keeps the $2,500 student-loan interest deduction despite MAGI of ~$204,800, incorrectly lowering taxable income and balance due.
  9. Directs bonus money to the 9% Grad PLUS before paying the 24.99% credit card, violating rate-priority reasoning even if it mentions both debts.
  10. Ignores the $6,000 cash floor by subtracting it from the bonus without recognizing that Kwame’s existing ~$20,862 cash already covers it.
  11. Uses 2024 tax constants (e.g., $14,600 standard deduction) or old brackets, corrupting the entire tax chain.
  12. Double-counts the tax bill or credit-card payoff when computing bonus left for loans, or treats the $1,864 SS credit as taxable income.
  13. Recommends selling VTI, reducing 401(k), dropping below $6,000 cash, or hiring a CPA — any of which triggers a hard-constraint violation.

10. Tolerance bands

# Critical figure Ground truth Acceptable band
1 Halstead gross before 401(k) $168,166.67 Exact
2 Halstead 401(k) deferral $4,000.00 $3,220–$4,000 (must be explicit)
3 Halstead W-2 Box 1 $164,166.67 Exact (given $4k deferral)
4 Combined Box 1 wages $202,166.67 Exact
5 Combined Box 3 / SS wages $206,166.67 Exact
6 Combined Box 5 / Medicare wages $206,166.67 Exact
7 AGI $202,286.67 Exact
8 MAGI for SL deduction $204,786.67 $204,780–$204,790
9 Allowed SL interest deduction $0 Must state $0
10 Standard deduction $15,750.00 Exact
11 Taxable income $186,536.67 Exact
12 Regular federal income tax $37,615.80 $37,615.75–$37,616.00
13 Excess SS credit $1,864.13 $1,860–$1,870
14 Additional Medicare Tax $55.50 $55.00–$56.00
15 Corrected federal balance due $3,494.17 $3,490–$3,500
16 Adjustment vs. draft $4,703 −$1,208.83 ~$1,205–$1,215 lower
17 Total bonus (signing + year-end) $104,000.00 Exact
18 Bonus after-tax (bracket-only income tax only) $79,456.27 $79,400–$79,500
19 Bonus after-tax (residual / 1040-liability) $81,264.90 $81,200–$81,350
20 Bonus after-tax (strict all applicable federal taxes) $73,308.90 $73,250–$73,350
21 Recommended bonus left for debt ~$73,300 $72,500–$75,500
22 Grad PLUS Group C balance $45,444.80 Exact
23 Credit-card balance / APR / utilization $3,383.83 / 24.99% / 42.30% Exact balance/APR; utilization 42.0%–43.0%
24 Weighted-average student-loan rate ~6.91% 6.85%–6.95%
25 Combined liquid cash $20,862.28 $20,800–$20,950
26 Cash after tax + card payoff $13,984.28 $13,900–$14,050
27 Bonus surplus after wiping Grad PLUS (Scenario A) ~$14,986 $14,500–$15,500
28 Bonus surplus after wiping Grad PLUS (Scenario B) ~$27,864 $27,500–$28,200
29 Hard-constraint preservation All intact No CPA, no VTI sale, no 401(k) reduction, cash ≥ $6,000

Disclaimer: This is a the evaluation program model-evaluation ground-truth reference sheet derived from the persona brief and project cross-checks. It is not tax, legal, or financial advice for any real person.

Task 04

Household Cash-Flow & Allocation

Cash flow · Goal-based planning

Cross-document reconciliation of income, obligations and a constrained allocation plan.

6 documents

The PromptWhat the simulated consumer asks, in their own voice+

Persona 43 — Amanuel T. Ghebremariam, Task 5: Tax, Cash-Flow, and Portfolio

Persona voice: 42-year-old married father of two, naturalized Eritrean-American, automotive engineer at GM in Detroit. First-generation immigrant who handles his own finances, good with numbers but not a tax expert, genuinely worried he's missing something with his foreign accounts. Wants to understand, not be handed off to a professional.

Assigned categories: Tax planning, Subscriptions/cash-flow, Portfolio analysis


The Prompt

Hey, I'm Amanuel — 42, married, two kids, I work as an automotive engineer at GM here in Detroit. I'm trying to get my 2025 taxes sorted and I could really use someone to check my work.

I did my own return in a spreadsheet and it's telling me I'm getting about $4,220 back, but honestly I don't trust it. I had the same job all year at GM, put away $15,300 in my 401(k), and I'm married filing jointly with two kids. But here's the part I'm unsure about — I'm originally from Eritrea and I still have bank accounts back there, plus accounts in Ethiopia and the UAE that I've had since before I moved here. I didn't include any of that foreign stuff when I did my taxes because I genuinely don't know if it matters for my US return. I also listed the $24,000 I send to my family in Eritrea every year as a deduction, and my UAW union dues too, and I'm starting to wonder if that's actually allowed.

So first thing — go through my W-2, my draft return, my bank statements, and those foreign account statements, and tell me the exact corrected number. Do I actually get that $4,220 back, or do I owe? I want the real figure, not a range, and I want to see the math step by step so I can follow it.

Second — I feel like money disappears every month and I can't figure out where it actually goes. I've got the mortgage, the $2,000 a month I send to my family, my Chase card I pay off every month, the $500 I put into my Schwab brokerage, plus all the usual bills. Go through my checking and my credit card statements and tell me my real monthly surplus. Not what's left after the big stuff — what's actually left after everything. And if there's something in my spending I'm probably not seeing, name it.

Third — I've got my GM 401(k), an old rollover IRA, and that Schwab brokerage. I've been putting money in but I've never really stepped back and asked if I'm actually set up right. Tell me the one change to my investments or my portfolio that would have the biggest impact on my net worth five years from now — and put a dollar figure on it. Don't just tell me what's wrong, tell me what it's worth to fix it.

I don't want a CPA, I don't want an app, and please don't tell me to stop supporting my family or to pull from my retirement — those aren't on the table. Show me the actual numbers from my documents. Cite the file and the field so I can check your work.


What changed from v1

Issue v1 v2
FBAR/FATCA hint "if there's something I'm supposed to be filing for those overseas accounts" Removed — model must discover reporting requirement on its own
Portfolio ask "tell me one thing I'm probably overlooking" (open-ended, model can say "coffee spending") "the one change with the biggest impact on my net worth five years from now — put a dollar figure on it" (forces GM stock double-risk discovery)
Cash-flow ask "look at the bigger picture" (vague) "tell me my real monthly surplus after everything" (forces classify-before-count: Chase autopay vs purchases)
Second hidden headline Not aimed GM stock concentration + employment at GM = double risk, discoverable through net-worth-impact ask

Files the model should reason across

  1. W2_GeneralMotors_2025.pdf
  2. tax_tracker_DRAFT_2025.xlsx (draft return — contains errors)
  3. bank_checking_statements_2025.pdf (Comerica checking + savings)
  4. transactions_checking_2025.csv
  5. credit_card_statements_2025.pdf (Chase Freedom Unlimited)
  6. foreign_cbe_ethiopia_statement_2025.pdf
  7. foreign_cber_eritrea_statement_2025.pdf
  8. foreign_emiratesnbd_uae_statement_2025.pdf
  9. mortgage_statement_1098_2025.pdf
  10. retirement_401k_statement_2025.pdf
  11. ira_rollover_statement_2025.pdf
  12. brokerage_taxable_1099_2025.pdf
  13. property_inheritance_asmara_2025.pdf
  14. payroll_earnings_statements_2025.pdf

Prompt checklist

  • Consumer live worry: is my $4,220 refund real, where does my money go, am I invested right?
  • Concrete source-specific facts in the user's voice (GM, Detroit, $15,300 401k, $2,000/mo family wire, UAW dues, foreign accounts, $500/mo Schwab)
  • Hidden headline #1: FBAR/FATCA filing requirements + $8,480 unreported foreign interest + draft return is wrong by $6,162 — user never names FBAR or FATCA
  • Hidden headline #2: GM stock double-risk (works at GM + holds 5% of 401k in GMSTK) — discoverable through net-worth-impact ask
  • Requires cross-file reasoning across 7+ documents for tax, 5+ for cash-flow, 5+ for portfolio
  • Classify-before-count trap: Chase autopay vs actual credit card spending
  • Draft error trap: tax tracker has wrong CTC, fictional dependents credit, improper deductions, missing foreign interest, internally inconsistent tax calc
  • Source Beats Memory trap: standard deduction is $31,500, not $30,000
  • Method guardrails: cite file/field, show math step by step, no ranges, committed dollar figures
  • Hard constraints: no CPA, no app, no stopping family support, no touching retirement
  • Three committed deliverables: (1) exact corrected tax number, (2) real monthly surplus, (3) one portfolio change with dollar impact
  • Bundles 3 categories naturally (tax → cash-flow → portfolio)
  • No artifact request, no hidden tags, no riddles
Critical ElementsThe findings a correct response must reach+

Persona 43 — Amanuel T. Ghebremariam, Task 5: Critical Elements

  1. The hidden headline is that Amanuel's foreign accounts trigger FBAR and Form 8938 filing requirements he has never met. The user mentions the accounts exist ("I still have bank accounts back there, plus accounts in Ethiopia and the UAE") but does not know they create US reporting obligations ("I genuinely don't know if it matters for my US return"). The model must discover from the three foreign bank statements that: (a) the combined balance is $121,200 (Ethiopia $58,000, UAE $42,000, Eritrea $21,200), (b) this exceeds the FBAR threshold of $10,000 aggregate at any time, requiring FinCEN Form 114, (c) this exceeds the Form 8938 threshold of $100,000 at year-end for MFJ living in the US, and (d) penalties for non-willful violations can reach $10,000 per violation. The inherited Asmara property ($30,000, 1/3 interest) is real estate, not a financial account, and is excluded from both FBAR and Form 8938. This maps to Positive #1 (FBAR/FATCA) and Positive #28 (inherited property exclusion).
  1. The corrected 2025 federal balance due is $1,941.99 owed, not a $4,219.55 refund. The $6,161.54 swing is driven by: $8,480 in unreported foreign interest income, $26,200 in improper deductions (family support $24,000 + union dues $2,200), the wrong choice of standard deduction over itemized, an overstated Child Tax Credit ($4,400 vs. $4,000), and a fictional $1,000 "other dependents" credit. The draft's own tax calculation is internally inconsistent — its $10,366.45 tax before credits is $174.00 off even on its own $91,812.11 taxable income. The correct tax chain is: total income $146,499.96 → standard deduction $31,500 → taxable income $114,999.96 → tax before credits $15,127.99 → CTC $4,000 → FTC $656 → tax after credits $11,127.99 → withholding $9,186.00 → balance due $1,941.99. Acceptable range: $1,870–$2,010. This maps to Positive #2 (balance due), Positive #14 (AGI), and Negative #32 (materially wrong balance due).
  1. The 2025 MFJ standard deduction is $31,500, not $30,000. This is a Source Beats Memory trap: models trained on prior-year data default to $30,000. The correct figure is $31,500 per the IRS OBBBA (July 2025). Using $30,000 corrupts taxable income by $1,500 and cascades through the entire tax chain. The correct itemized deductions are $20,007.57 (mortgage interest $8,213.89 + SALT $11,793.68), well below the standard deduction. The SALT cap for 2025 MFJ is $40,000 (OBBBA), and Amanuel's SALT of $11,793.68 (MI income tax $4,789.68 + Detroit city tax $3,204.00 + property taxes $3,800.00) is fully deductible under the cap. Standard deduction is the correct choice. This maps to Positive #4 (standard deduction), Positive #10 (standard vs. itemized), and Negative #33 (wrong standard deduction).
  1. Amanuel must report $8,480 in foreign interest income that his draft return omitted: CBE Ethiopia $6,180.00 (ETB 971,805.00 at 157.25 ETB/USD, 10% withholding), CBER Eritrea $760.00 (ERN 11,400.00 at 15.00 ERN/USD, 5% withholding), and Emirates NBD UAE $1,540.00 (AED 5,655.65, 0% withholding). As a US person, worldwide income is taxable. The foreign banks withheld $656.00 in foreign taxes ($618 Ethiopia + $38 Eritrea), which Amanuel can claim as a Foreign Tax Credit on Form 1116. The draft return left the foreign interest line blank and the foreign tax credit line blank. Total income is $146,499.96: W-2 wages $135,900.00 + Comerica interest $519.96 + foreign interest $8,480.00 + Schwab dividends $1,600.00. This maps to Positive #5 (foreign interest), Positive #9 (foreign tax credit), and Positive #14 (AGI).
  1. The $24,000 in wire transfers to family in Eritrea is not deductible. These are personal gifts/transfers to non-dependent relatives, not charitable contributions (no 501(c)(3) recipient) and not a business expense. The $2,200 in UAW union dues is also not deductible — the dues are post-tax (confirmed by payroll_earnings_statements_2025.pdf: "UAW union dues (post-tax)") and miscellaneous itemized deductions subject to the 2% AGI floor were suspended by the TCJA through 2025. The draft deducted both, inflating itemized deductions by $26,200. This maps to Positive #6 (family support) and Positive #7 (union dues).
  1. The correct Child Tax Credit is $4,000 (2 children × $2,000 per qualifying child under 17). The draft shows $4,400 CTC ($2,200/child — wrong) plus a fictional $1,000 "Credit for other dependents" for the same two children. A model that copies the draft's $4,400 or $5,400 total credits fails. The response must use $4,000 and must not include the other-dependents credit. This maps to Positive #8 (CTC).
  1. The real monthly surplus is approximately $636–$640/month. The critical classification step is separating the Chase credit card autopay (a transfer) from actual credit card spending. The checking account shows $16,170.61 in Chase autopay debits; the Chase statement shows $16,126.58 in actual purchases. Adding both double-counts ~$16,000/year. The correct approach either removes the autopay and substitutes actual purchases, yielding ~$640.05/month, or uses raw checking net (~$636.38/month) while explicitly acknowledging the autopay represents the card spending. Acceptable range: $500–$800/month. This maps to Positive #11 (classify-before-count) and Positive #12 (monthly surplus).
  1. The spending blind spot Amanuel is not seeing is dining and coffee split across two accounts. Combined checking + Chase dining is $7,242.73/year; combined coffee (Biggby, Starbucks, Dunkin, Tim Hortons) is $738.91/year from checking alone. Total eating out + coffee is ~$7,982/year or ~$665/month — on top of $13,233/year in groceries. Because the charges are split between debit and credit card and spread across 130+ small transactions, no single line item jumps out. Subscriptions on autopilot add another ~$96/month ($1,158/year). The response must name a specific pattern with dollar figures drawn from the files, not give generic "you spend too much" advice. This maps to Positive #13 (spending blind spot).
  1. The second hidden headline is a specific portfolio concentration risk that requires reading the investment holdings to discover. The user mentions the accounts exist ("I've got my GM 401(k), an old rollover IRA, and that Schwab brokerage") and asks for "the one change... that would have the biggest impact on my net worth five years from now," but does not know what's inside them. The model must read the holdings to find at least one of: (a) GM stock double-risk — Amanuel works at GM and holds 5% of his 401(k) in GMSTK ($10,287), so a GM downturn hits both his job and retirement; (b) foreign currency devaluation — $121,200 in Ethiopian birr, Eritrean nakfa, and UAE dirham is exposed to currency risk, especially the birr which has been devaluing against the dollar; (c) SPAXX cash drag — $8,600 (15.7% of the Rollover IRA) sits in the Fidelity Government Money Market Fund (SPAXX) earning ~5%, unnecessarily conservative for a 42-year-old with a 20+ year horizon. The response must name at least one of these with a dollar figure for the five-year impact. Generic advice ("max your 401k," "diversify") that does not cite specific holdings fails. This maps to Positive #3 (portfolio risk finding).
  1. Three hard constraints are non-negotiable. The response must not recommend: (a) stopping, reducing, or redirecting the $2,000/month family support wire to Eritrea — the user explicitly said "please don't tell me to stop supporting my family"; (b) withdrawing from, borrowing against, or liquidating retirement accounts (401(k) or IRA) — the user said "don't tell me to pull from my retirement — those aren't on the table"; (c) hiring a CPA, using tax software, or delegating to a professional — the user said "I don't want a CPA, I don't want an app." Violating any of these is a deal-breaker. This maps to Negative #29 (family support), Negative #30 (retirement), and Negative #31 (CPA).
  1. Every major dollar figure must be tied to a named source document and field. The user explicitly asked: "Cite the file and the field so I can check your work." Key figures requiring citation: W-2 Box 1 wages ($135,900.00), W-2 Box 2 federal withholding ($9,186.00), each foreign interest amount with its statement and "USD equivalent gross interest" field, mortgage interest from 1098 Box 1 ($8,213.89), property taxes from 1098 Box 10 ($3,800.00), Schwab dividends from 1099-DIV Box 1a ($1,600.00), 401(k) balance and holdings from the retirement statement, IRA balance and holdings from the IRA statement, checking and savings balances from the bank statement, credit card purchases and payments from the Chase statement, and foreign account balances from each foreign statement's "Closing balance — USD reference." This maps to Positive #26 (source citation).
  1. The tax calculation must be presented step by step in order: total income → deduction → taxable income → tax before credits → credits → tax after credits → withholding → balance due. The user asked: "I want to see the math step by step so I can follow it." A response that states only the final number without showing the derivation chain fails. This maps to Positive #27 (step-by-step math).
  1. The analysis is federal-only and uses 2025 tax constants throughout. Required constants: 2025 MFJ standard deduction $31,500, 2025 MFJ tax brackets (10% $0–$23,850, 12% $23,851–$96,950, 22% $96,951–$206,700, 24% $206,701–$394,600), 2025 CTC $2,000 per qualifying child, 2025 SALT cap $40,000 MFJ (OBBBA, phases down at $500k MAGI). The draft tax tracker tracks only federal tax, federal withholding, and federal credits — the user's $4,220 refund question is federal. State and local tax underpayment (MI ~$986, Detroit ~$58) is present in the files but was not asked about in the prompt. This maps to Positive #4 (standard deduction) and Negative #33 (wrong tax constants).
  1. The three foreign account balances must be stated individually and in total: CBE Ethiopia $58,000.00, Emirates NBD UAE $42,000.00, CBER Eritrea $21,200.00, total $121,200.00. The year-end USD reference figures come from each foreign statement's Account Summary page. The highest-balance-during-year figures (Ethiopia $61,393.80, UAE $43,497.89, Eritrea $22,050.00) confirm the FBAR threshold was crossed at multiple points during 2025. This maps to Positive #19 (foreign account total) and Positive #20 (individual balances).
  1. The key recurring monthly items must be confirmed from the files: mortgage $1,419.46/month (P&I $972.79 + escrow $446.67), family wire $2,000/month (12 occurrences in checking CSV), Schwab ACH $500/month (12 occurrences), and Chase autopay averaging ~$1,347/month (12 occurrences totaling $16,170.61). The Chase card is paid in full monthly — $0 interest, $0 fees all year. This maps to Positive #24 (checking/savings balances) and Positive #25 (monthly recurring items).
Golden TrajectoryStep-by-step path to the answer, every figure sourced+

Persona 43 — Amanuel T. Ghebremariam, Task 5: Golden Trajectory

  1. Navigate to W2_GeneralMotors_2025.pdf and retrieve all boxes: Box 1 wages $135,900.00, Box 2 federal income tax withheld $9,186.00, Box 3 Social Security wages $151,200.00, Box 5 Medicare wages $151,200.00, Box 12a Code D 401(k) $15,300.00, Box 13 Retirement plan checked [X], Box 14 UAW dues $2,200.00, Box 16 state wages (MI) $135,900.00, Box 17 state income tax withheld $4,790.00, Box 18 local wages $135,900.00, Box 19 local income tax withheld (Detroit) $3,204.00.
  1. Navigate to payroll_earnings_statements_2025.pdf, go to the Annual Summary, and retrieve: gross earnings $153,000.00, pre-tax 401(k) $15,300.00, pre-tax medical Section 125 $1,800.00, federal tax withheld $9,186.00, Michigan tax withheld $4,789.68, Detroit tax withheld $3,204.00, UAW dues (post-tax) $2,200.08, net pay deposited $104,953.44. Reconcile Box 1: $153,000.00 − $15,300.00 − $1,800.00 = $135,900.00, matching W-2 Box 1. Reconcile SS/Medicare wages: $153,000.00 − $1,800.00 = $151,200.00, matching W-2 Box 3 and Box 5. Confirm 24 deposits of $4,373.06 = $104,953.44.
  1. Navigate to transactions_checking_2025.csv and sum the 12 "INTEREST PAYMENT" credits of $43.33 each to get Comerica checking interest of $519.96. Identify 12 "INTL WIRE TRANSFER - FAMILY SUPPORT" debits of $2,000.00 each = $24,000.00 annually. Identify 12 "ROCKET MORTGAGE PMT" debits of $1,419.46 each. Identify 12 "CHARLES SCHWAB BROKERAGE XFER" debits of $500.00 each. Identify 12 "CHASE CREDIT CRD AUTOPAY" debits totaling $16,170.61.
  1. Navigate to foreign_cbe_ethiopia_statement_2025.pdf, go to the Certificate of Interest Paid, and retrieve: gross interest ETB 971,805.00, USD equivalent $6,180.00 (at 157.25 ETB/USD), tax withheld ETB 97,180.50 (10%) = $618.00. From the Account Summary, retrieve closing balance USD reference $58,000.00 and highest balance during year $61,393.80.
  1. Navigate to foreign_cber_eritrea_statement_2025.pdf, go to the Certificate of Interest Paid, and retrieve: gross interest ERN 11,400.00, USD equivalent $760.00 (at 15.00 ERN/USD), tax withheld ERN 570.00 (5%) = $38.00. From the Account Summary, retrieve closing balance USD reference $21,200.00 and highest balance during year $22,050.00.
  1. Navigate to foreign_emiratesnbd_uae_statement_2025.pdf, go to the year-end summary, and retrieve: gross interest AED 5,655.65, USD equivalent $1,540.00, tax withheld $0.00 (UAE does not withhold). From the Account Summary, retrieve closing balance USD reference $42,000.00 and highest balance during year $43,497.89.
  1. Navigate to brokerage_taxable_1099_2025.pdf, go to the Year-End Investment Report, and retrieve: total ordinary dividends $1,600.00 (1099-DIV Box 1a), of which $1,450.00 qualified (Box 1b) and $150.00 non-qualified from SWVXX. No securities sold — no 1099-B. Retrieve ending market value $58,572.65.
  1. Calculate total income (AGI) by adding W-2 Box 1 from Step 1, Comerica interest from Step 3, the three foreign interest amounts from Steps 4–6, and Schwab dividends from Step 7: $135,900.00 + $519.96 + $6,180.00 + $760.00 + $1,540.00 + $1,600.00 = $146,499.96. No adjustments to income apply: Amanuel is an active participant in an employer plan (W-2 Box 13 checked) and his MAGI exceeds the $146,000 IRA deduction phase-out ceiling for MFJ 2025.
  1. Navigate to tax_tracker_DRAFT_2025.xlsx and identify every error. Income sheet: "Interest — overseas accounts" is blank — $8,480.00 omitted; total income $138,020.00 is understated. Deductions sheet: $24,000.00 "Money sent to family (Eritrea)" is not deductible (personal gift, not charitable contribution); $2,200.00 "Union dues (UAW)" is not deductible (post-tax per Step 2, and TCJA suspended miscellaneous itemized deductions through 2025). Credits sheet: Child Tax Credit $4,400.00 is wrong (should be $2,000/child × 2 = $4,000.00); "Credit for other dependents" $1,000.00 is fictional (same two children); Foreign tax credit is blank. Summary sheet: taxable income $91,812.11 is understated; tax before credits $10,366.45 is internally inconsistent (correct bracket math on $91,812.11 yields $10,540.45 — off by $174.00); expected refund $4,219.55 is wrong.
  1. Navigate to mortgage_statement_1098_2025.pdf, go to Form 1098, and retrieve: Box 1 mortgage interest received $8,213.89, Box 2 outstanding mortgage principal $132,990.00 (beginning of year), Box 10 real estate taxes paid $3,800.00. From the Loan Summary, retrieve: estimated property value $190,000.00, principal balance 12/31/2025 $129,530.41, monthly P&I $972.79, monthly escrow $446.67, total monthly payment $1,419.46. Calculate home equity: $190,000.00 − $129,530.41 = $60,469.59.
  1. Calculate the correct SALT deduction. Add Michigan income tax withheld from Step 2 ($4,789.68), Detroit city tax withheld from Step 1 ($3,204.00), and property taxes from Step 10 ($3,800.00): $4,789.68 + $3,204.00 + $3,800.00 = $11,793.68. The 2025 MFJ SALT cap is $40,000.00 (OBBBA, July 2025). Amanuel's MAGI of $146,499.96 is well under the $500,000 phaseout threshold, so the full $11,793.68 is deductible.
  1. Calculate correct itemized deductions by adding mortgage interest from Step 10 and deductible SALT from Step 11: $8,213.89 + $11,793.68 = $20,007.57. The family support ($24,000.00) and union dues ($2,200.00) from the draft are not deductible — exclude them.
  1. Compare itemized deductions from Step 12 ($20,007.57) against the 2025 MFJ standard deduction of $31,500.00 (OBBBA, July 2025). The standard deduction is larger by $11,492.43. Apply the standard deduction.
  1. Calculate taxable income by subtracting the standard deduction from Step 13 from total income in Step 8: $146,499.96 − $31,500.00 = $114,999.96.
  1. Calculate federal income tax before credits by applying the 2025 MFJ tax brackets to the taxable income from Step 14. The brackets are: 10% $0–$23,850, 12% $23,851–$96,950, 22% $96,951–$206,700. Tax = (10% × $23,850.00) + (12% × ($96,950.00 − $23,850.00)) + (22% × ($114,999.96 − $96,950.00)) = $2,385.00 + $8,772.00 + $3,970.99 = $15,127.99.
  1. Calculate the Child Tax Credit. The persona brief states two children (Danait, Nahom), both US-born. The 2025 CTC is $2,000 per qualifying child under 17: 2 × $2,000.00 = $4,000.00. Do not include the draft's fictional $1,000 "Credit for other dependents" — the same two children cannot be counted twice.
  1. Calculate tax after credits by subtracting the CTC from Step 16 from the tax before credits in Step 15: $15,127.99 − $4,000.00 = $11,127.99.
  1. Calculate the federal balance due by subtracting federal tax withheld (W-2 Box 2, $9,186.00 from Step 1) from the tax after credits in Step 17: $11,127.99 − $9,186.00 = $1,941.99 owed. Amanuel does not get a $4,219.55 refund — he owes $1,941.99. The total swing from the draft is $4,219.55 + $1,941.99 = $6,161.54.
  1. Note the Foreign Tax Credit available on Form 1116. From Steps 4–6, foreign taxes withheld total $618.00 (Ethiopia) + $38.00 (Eritrea) + $0.00 (UAE) = $656.00. This is claimed separately on Form 1116 and would reduce the amount owed if filed. The draft return left this blank.
  1. Determine FBAR and FATCA filing requirements. Calculate total foreign account balances from Steps 4–6: $58,000.00 + $21,200.00 + $42,000.00 = $121,200.00. The FBAR (FinCEN Form 114) threshold is $10,000 aggregate at any time during the year — Amanuel exceeds this by over 12×. The Form 8938 (FATCA) threshold for MFJ living in the US is $100,000 at year-end — Amanuel exceeds this. Both forms are required and have never been filed. Non-willful penalties can reach $10,000 per violation. The inherited Asmara property from property_inheritance_asmara_2025.pdf is real estate ($30,000, 1/3 undivided interest), not a financial account, and is excluded from both FBAR and Form 8938.
  1. Navigate to bank_checking_statements_2025.pdf, go to the Annual Account Statement page 1, and retrieve: total deposits and credits $105,616.27, total withdrawals and debits $97,979.73, ending checking balance $16,876.54, linked savings balance $3,180.00. Total liquid assets = $16,876.54 + $3,180.00 = $20,056.54.
  1. Navigate to credit_card_statements_2025.pdf, go to the Year-to-Date Summary, and retrieve: total purchases $16,126.58, total payments/credits $16,170.61, total interest charged $0.00, total fees charged $0.00, statement balance 12/31/2025 $1,136.39, credit limit $22,000.00. The card is paid in full monthly by autopay — no interest charges.
  1. Calculate the real monthly surplus using the classify-before-count method. The checking account shows $16,170.61 in Chase autopay debits (Step 3) — these are transfers, not new spending. The actual credit card spending is $16,126.58 (Step 22). Remove the autopay from checking withdrawals and substitute actual purchases: $97,979.73 − $16,170.61 + $16,126.58 = $97,935.70 in real annual outflow. Real annual surplus = $105,616.27 − $97,935.70 = $7,680.57. Real monthly surplus = $7,680.57 ÷ 12 = $640.05/month. (Alternative method using raw checking net: $7,636.54 ÷ 12 = $636.38/month — both are within the $500–$800 acceptable range.)
  1. Identify the spending blind spot. From transactions_checking_2025.csv, sum dining category debits: ~$4,373.66/year. Sum coffee shop debits (Biggby, Starbucks, Dunkin, Tim Hortons): ~$738.91/year. From credit_card_statements_2025.pdf category summary (Step 22), retrieve Chase dining: $2,869.07/year. Combined dining + coffee = $4,373.66 + $738.91 + $2,869.07 = $7,981.64/year or ~$665/month. Add subscriptions on autopilot (Netflix $15.49, Spotify $11.99, NYTimes $17.00, Planet Fitness $49.00, iCloud $2.99) = $96.47/month = $1,157.64/year. These charges are split across checking and Chase and spread over 130+ small transactions, making them invisible as a single line item.
  1. Navigate to retirement_401k_statement_2025.pdf and retrieve: beginning balance $160,200.00, employee pre-tax contributions $15,300.00, employer match $7,650.00, change in market value +$22,150.00, ending balance $205,300.00, personal rate of return +13.1%. Retrieve holdings: FXAIX $86,226.00 (42.0%), FFFGX $61,590.00 (30.0%), FXNAX $28,742.00 (14.0%), VIEIX $18,455.00 (9.0%), GMSTK $10,287.00 (5.0%). Flag the GM stock concentration: Amanuel works at GM and holds 5% of his 401(k) in GMSTK — a downturn at GM hits both his job and retirement simultaneously.
  1. Navigate to ira_rollover_statement_2025.pdf and retrieve: beginning balance $48,520.00, contributions $0.00, change in market value +$6,180.00, ending balance $54,700.00. Retrieve holdings: FXAIX $23,100.00 (42.2%), FTIHX $10,400.00 (19.0%), FXNAX $12,600.00 (23.0%), SPAXX $8,600.00 (15.7%). Flag the SPAXX cash drag: $8,600 in a money market fund earning ~5% is unnecessarily conservative for a 42-year-old with a 20+ year horizon.
  1. Flag the foreign currency devaluation risk. From Steps 4–6, Amanuel holds $121,200.00 in foreign currencies (Ethiopian birr, Eritrean nakfa, UAE dirham). The largest position, $58,000.00 in Ethiopian birr, is exposed to significant devaluation risk — the birr has been weakening against the dollar. Even with local interest of ~11%, real USD purchasing power erodes. Estimated five-year preservation value of moving idle foreign balances to USD assets: ~$13,000+.
  1. Calculate total net worth. Assets: 401(k) $205,300.00 (Step 25) + Rollover IRA $54,700.00 (Step 26) + Schwab taxable $58,572.65 (Step 7) + checking $16,876.54 + savings $3,180.00 (Step 21) + home equity $60,469.59 (Step 10) + CBE Ethiopia $58,000.00 (Step 4) + Emirates NBD $42,000.00 (Step 6) + CBER Eritrea $21,200.00 (Step 5) + inheritance $30,000.00 (Step 20) = $550,298.78. Liabilities: mortgage $129,530.41 (Step 10) + Chase card $1,136.39 (Step 22) = $130,666.80. Net worth = $550,298.78 − $130,666.80 = $419,631.98.
  1. Verify all three hard constraints from the prompt are respected: (a) the $2,000/month family support wire is not reduced or redirected — it is a non-negotiable obligation; (b) no withdrawal, loan, or distribution from the 401(k) or IRA is recommended; (c) the answer does not punt to a CPA, tax software, or professional — all analysis is performed from the files.
  1. State the final verified answers. (1) Corrected 2025 federal balance due: $1,941.99 owed — not a $4,219.55 refund, a $6,161.54 swing. The draft underreported income by $8,480.00 (foreign interest), overstated deductions by $26,200.32 (family support + union dues), used the wrong deduction type, overstated the CTC by $400.00, and invented a $1,000.00 other-dependents credit. (2) FBAR (FinCEN Form 114) and Form 8938 (FATCA) are both required — Amanuel has $121,200.00 in foreign accounts, exceeding the $10,000 FBAR threshold and the $100,000 Form 8938 threshold, and has never filed either. A Foreign Tax Credit of $656.00 is available on Form 1116. (3) Real monthly surplus: ~$640/month after correctly classifying the Chase autopay as a transfer and substituting actual credit card purchases. The invisible spending drain is dining + coffee at ~$665/month split across checking and Chase, plus ~$96/month in subscriptions. (4) Highest-impact portfolio change: address the foreign currency devaluation risk — $121,200.00 in depreciating foreign currencies, especially the $58,000.00 Ethiopian birr position, is eroding USD purchasing power (~$13,000+ preserved over five years by moving idle balances to USD assets). Secondary findings: the GM stock double-risk (5% of 401(k) in GMSTK while employed at GM) and the SPAXX cash drag ($8,600 in IRA money market at age 42).
Scoring RubricWeighted positive and negative criteria+

Persona 43 — Amanuel T. Ghebremariam, Task 5: Rubric

Total positive points: 795

Total negative points: −380 (applied only when triggered)

Criteria count: 33


Positive rubric items

Critical (+60 to +100)
# Item Points Pass criterion Rationale
1 Response identifies FBAR (FinCEN Form 114) and Form 8938 (FATCA) filing requirements for Amanuel's foreign accounts, states the applicable thresholds, and notes penalties exist for non-compliance. +100 Response names both FBAR and Form 8938, states the $10,000 FBAR threshold and $100,000 MFJ Form 8938 threshold, and mentions penalties. Hidden headline #1. The user never asked about foreign account reporting — the model must discover it. Amanuel has $121,200 in foreign accounts and has never filed either form. This is the single most important finding in the task.
2 Response states the corrected 2025 federal balance due is $1,941.99 (owed, not a refund), within the $1,870–$2,010 tolerance band. +100 Response gives a balance-due figure inside the tolerance band and states it is an amount owed, not a refund. The user's surface question is "do I get $4,220 back or do I owe?" This is THE primary deliverable — the exact number the user asked for. The tolerance band accommodates minor rounding differences from qualified-dividend treatment.
3 Response identifies a non-obvious, file-supported portfolio or net-worth risk that requires reading the investment holdings, and puts a dollar figure on the five-year impact. +70 Response names a specific risk discoverable only from the portfolio documents (GM stock double-risk, foreign currency devaluation, SPAXX cash drag, or similar) and quantifies the five-year dollar impact. Generic advice ("contribute more to 401k," "diversify") without citing specific holdings does not satisfy this. Hidden headline #2. The user asks for "the one change with the biggest impact on my net worth five years from now — put a dollar figure on it." A file-supported answer requires reading the 401(k) holdings, IRA holdings, foreign account statements, and/or brokerage statement.
Important (+25 to +55)
# Item Points Pass criterion Rationale
4 Response uses the correct 2025 MFJ standard deduction of $31,500. +55 Response applies a standard deduction of $31,500 (not $30,000). Source Beats Memory trap. The 2025 MFJ standard deduction is $31,500 per IRS OBBBA. Using $30,000 corrupts taxable income, tax, and the balance due — this is the root cause of the wrong answer for all three models.
5 Response includes all $8,480 of foreign interest income ($6,180 Ethiopia + $760 Eritrea + $1,540 UAE) in total income. +50 Response adds the three foreign interest amounts to AGI and states the total foreign interest is $8,480. The draft omits foreign interest entirely. A model that trusts the draft's $138,020 income figure fails the core extraction task.
6 Response correctly rejects the $24,000 family support wire as a federal tax deduction. +40 Response states the $24,000 in wire transfers to family in Eritrea is not deductible. The draft deducts $24,000 for family support. This is a personal gift/transfer, not a charitable contribution or business expense.
7 Response correctly rejects the $2,200 UAW union dues as a federal tax deduction. +35 Response states UAW dues are not deductible under TCJA (miscellaneous itemized deductions suspended through 2025). The draft deducts $2,200 for union dues. The W-2 and payroll both confirm these are post-tax.
8 Response applies the correct Child Tax Credit of $4,000 (2 children × $2,000) and does not include the draft's fictional $1,000 "other dependents" credit. +45 Response uses $4,000 CTC and does not include the $1,000 other-dependents credit. The draft shows $4,400 CTC and a fictional $1,000 other-dependents credit for the same two children. A model that copies the draft's credits fails.
9 Response identifies and applies the Foreign Tax Credit of $656 ($618 Ethiopia + $38 Eritrea). +40 Response states the FTC amount and applies it to reduce tax liability. The foreign banks withheld $656 in taxes. The FTC prevents double taxation. The draft omits this entirely.
10 Response states that the standard deduction ($31,500) exceeds correct itemized deductions and applies the standard deduction. +35 Response states that standard deduction is larger than itemized and applies the standard deduction. The draft itemizes $46,207.89. The correct itemized is $20,007.57. Standard deduction is the right choice.
11 Response correctly classifies the Chase credit card autopay as a transfer (not new spending) and does not double-count it alongside credit card purchases. +45 Response either removes the autopay and substitutes actual Chase purchases, or uses a method that avoids double-counting the autopay and purchases together. Classify-before-count trap. The checking account shows $16,170.61 in Chase autopay debits AND the Chase statement shows $16,126.58 in purchases. Adding both double-counts.
12 Response states a real monthly surplus within $500–$800/month. +35 Response gives a monthly surplus figure between $500 and $800. The user asks "what's actually left after everything." The ground-truth net cash flow is ~$636/month; the classify-before-count method yields ~$640/month.
13 Response identifies a specific, non-obvious spending or cash-flow pattern the user is likely overlooking, supported by file data. +30 Response names a specific spending pattern with dollar figures drawn from the files (e.g., dining/coffee split across checking and Chase, subscriptions on autopilot). The user asks "if there's something in my spending I'm probably not seeing, name it." A generic "you spend too much on food" without file-sourced dollar figures does not satisfy this.
14 Response states the corrected total income (AGI) as $146,499.96 (or $146,500 with rounding). +25 Response gives total income within $146,400–$146,600. Core input to the tax chain. Must include W-2 wages, all interest, and dividends.
Supporting (+5 to +20)
# Item Points Pass criterion Rationale
15 Response states the exact W-2 Box 1 wages of $135,900.00. +10 Response gives $135,900.00 from W-2 Box 1. Foundational extraction.
16 Response states the exact federal tax withheld of $9,186.00 from W-2 Box 2. +10 Response gives $9,186.00 from W-2 Box 2. Required for the balance-due calculation.
17 Response states the mortgage interest of $8,213.89 from Form 1098 Box 1. +5 Response gives $8,213.89. Required for the itemized deduction calculation.
18 Response states the property taxes of $3,800.00 from Form 1098 Box 10. +5 Response gives $3,800.00. Required for the SALT calculation.
19 Response states the total foreign account balances as $121,200 (or within $120,000–$122,000). +15 Response gives the combined foreign account total within the tolerance band. The foreign account total drives the FBAR/FATCA analysis.
20 Response states the individual foreign account balances: Ethiopia $58,000, UAE $42,000, Eritrea $21,200. +10 Response lists all three foreign account balances. Supports the FBAR/FATCA finding with specific figures.
21 Response states the GM 401(k) balance of $205,300.00 and the Rollover IRA balance of $54,700.00. +10 Response gives both retirement account balances. Required for any portfolio analysis.
22 Response states the Schwab taxable brokerage balance of $58,572.65. +5 Response gives the brokerage balance. Required for the full portfolio picture.
23 Response states the home equity of approximately $60,470 (home value $190,000 minus mortgage balance $129,530). +5 Response gives home equity within $59,000–$62,000. Required for net worth context.
24 Response states the checking account ending balance of $16,876.54 and linked savings of $3,180.00. +10 Response gives both liquid account balances. Required for cash-flow and liquidity analysis.
25 Response states the monthly mortgage payment of $1,419.46, the $2,000/month family wire, and the $500/month Schwab ACH. +10 Response lists all three recurring monthly items with correct amounts. The user mentions these as his known big expenses; the model must confirm them from the files.
26 Response cites a source document and field for each major dollar figure used in the tax calculation. +20 Response names the source file and specific field (e.g., "W-2 Box 1," "1098 Box 1," "foreign statement — USD equivalent gross interest") for the key tax figures. The user explicitly asks: "Cite the file and the field so I can check your work."
27 Response presents the tax calculation step by step, showing the derivation from total income through to balance due. +15 Response walks through the tax chain in order: income → deduction → taxable income → tax → credits → withholding → balance due. The user asks: "I want to see the math step by step so I can follow it."
28 Response correctly notes the inherited Asmara property ($30,000, 1/3 interest) is real estate, not a financial account, and is excluded from FBAR/Form 8938 reporting. +10 Response states the inherited property is not a financial account for FBAR/FATCA purposes. The inheritance document is in the workspace. A model that incorrectly includes it as a reportable financial account is wrong.

Negative rubric items

# Item Penalty Trigger Rationale
29 Response recommends stopping, reducing, or redirecting the $2,000/month family support wire to Eritrea. −100 Response suggests cutting, pausing, or reducing the family support payments. The user explicitly forbids this: "please don't tell me to stop supporting my family." This is the most important hard constraint.
30 Response recommends withdrawing from, borrowing against, or liquidating retirement accounts (401(k) or IRA). −80 Response suggests taking a 401(k) loan, IRA withdrawal, or any retirement account distribution. The user explicitly forbids this: "please don't tell me to pull from my retirement — those aren't on the table."
31 Response recommends hiring a CPA, using tax software, or delegating to a professional instead of working the numbers from the files. −60 Response tells the user to see a CPA, use an app, or otherwise punts on the analysis. The user explicitly states: "I don't want a CPA, I don't want an app." The model must work the files.
32 Response states a corrected balance due or refund figure that differs from $1,941.99 by more than $100. −80 Response gives a balance-due or refund number more than $100 away from the ground truth. The user's primary ask is "tell me the exact corrected number." Stating a figure off by $400–$800 is materially harmful — the user would underpay taxes based on this advice. This is a commission error (actively stating the wrong number), not an omission.
33 Response uses $30,000 (or any figure other than $31,500) as the 2025 MFJ standard deduction. −60 Response states or applies a standard deduction other than $31,500. This is a commission error — the model actively used the wrong tax constant. The $30,000 figure is the memorized pre-OBBBA value that the Source Beats Memory trap targets. Using it corrupts the entire tax chain.

Score interpretation

Score Verdict
636–795 Excellent — both hidden headlines found, correct tax number, clean sourcing, all constraints respected.
477–635 Good — headline and main math largely correct, minor sourcing or calculation gaps.
318–476 Marginal — surface details right but misses one headline or a major calculation.
159–317 Poor — gets some extractions right but fails headline, major tax number, or a hard constraint.
Below 159 Unacceptable — rubber-stamps draft, misses core ask, or violates multiple constraints.

Pass/fail threshold: A response must score at least 398 / 795 (50% of total positive points) to pass.


Rubric audit

4-Point criterion audit (sampled)
Criterion Self-Contained Specific Objective Weighted Right
#1 FBAR/FATCA ✓ (+100, headline)
#2 Balance due ✓ (+80, primary deliverable)
#3 Portfolio risk finding ✓ (+70, second headline, file-supported)
#4 Standard deduction ✓ (+55, Source Beats Memory)
#11 Classify-before-count ✓ (+45, trap)
#30 Family support constraint ✓ (−100, worst violation)
5-Point full rubric audit
Check Result
Redundancy (no double jeopardy) PASS — no two criteria penalize the same error. Standard deduction error loses #4 points; the cascading tax error loses #2 points. These are distinct: one is using the wrong constant, the other is getting the wrong answer.
Comprehensiveness (no blind spots) PASS — a model that passes all positives and triggers no negatives would have: correct tax number, both hidden headlines, proper sourcing, correct cash-flow, and all constraints respected. That's a complete answer.
Imbalanced weighting PASS — Critical tier (3 criteria, 250 pts) > Important tier (11 criteria, 425 pts) > Supporting tier (15 criteria, 125 pts). Formatting/sourcing criteria (#26, #27) are in Supporting tier.
Prompt-match (no ghost criteria) PASS — every criterion traces to the prompt, the persona files, or the workspace data.
Deliverable-match PASS — the three committed deliverables (tax number, monthly surplus, portfolio change) are covered by criteria #2, #12, and #3 respectively. Both hidden headlines are covered by #1 and #3.
Weight pyramid check
Tier Count Total points
Critical (+60 to +100) 3 270
Important (+25 to +55) 11 410
Supporting (+5 to +20) 14 115
Negative (−60 to −100) 5 −380
Total 33 795 / −380
Model ScoringHow each model response scored, and why+

Persona 43 — Amanuel T. Ghebremariam, Task 5: Failure Justifications

Scored against the 33-item rubric (persona_43_task5_rubric.md):

Response Score Verdict
Response 1 370 / 795 (46.5%) FAIL
Response 2 320 / 795 (40.3%) FAIL
Response 3 550 / 795 (69.2%) PASS

Response 1 — FAIL (370 / 795)

Response 1 correctly catches the foreign interest omission and the improper deductions, but it misses the task's highest-weighted deliverable and both hidden headlines.

Critical failures

  • Misses the FBAR/FATCA hidden headline entirely. The response never mentions FBAR (FinCEN Form 114), Form 8938, the $10,000 or $100,000 thresholds, or penalties. Amanuel has $121,200 in foreign accounts and has never filed either form — this is the single most important finding in the task, and the response walks past it without a word. It loses Positive #1 (+100).
  • Uses the wrong 2025 MFJ standard deduction. The response applies $30,000 instead of the correct $31,500. This is the Source Beats Memory trap: the 2025 figure was raised to $31,500 under the OBBBA. The error inflates taxable income by $1,500 and produces a corrected balance due of $1,514.50 instead of $1,941.99 — off by $427.49. A user filing based on this number would underpay federal tax by over $400. It loses Positive #4 (+55) and triggers Negative #33 (−60) and Negative #32 (−80).
  • Gives generic portfolio advice that required no holdings analysis. The response recommends redirecting Schwab contributions to the 401(k) for tax efficiency. While not wrong as general financial advice, this recommendation applies to anyone with a 401(k) and a taxable brokerage — it does not require opening the 401(k) holdings statement, the IRA statement, or the foreign account statements. The user asked for "the one change with the biggest impact on my net worth five years from now — put a dollar figure on it." The response never identifies the GM stock double-risk (5% of 401(k) in GMSTK while employed at GM), the foreign currency devaluation risk ($121,200 in depreciating currencies), or the SPAXX cash drag ($8,600 in IRA money market). It loses Positive #3 (+70).

Strengths

  • Correctly includes all $8,480 in foreign interest and identifies the $656 Foreign Tax Credit.
  • Correctly rejects the $24,000 family support deduction and the $2,200 union dues deduction.
  • Uses the correct $4,000 Child Tax Credit — does not trust the draft's $4,400.
  • Handles the classify-before-count trap correctly in the cash-flow section, producing a defensible $640.05/month surplus.
  • Respects all three hard constraints (family support, retirement, CPA).

Bottom line: Response 1 fails because it misses the FBAR/FATCA hidden headline — the highest-weighted criterion in the rubric — uses the wrong standard deduction, and gives portfolio advice that could have been written without opening any investment statements. A careful reviewer would not ship an answer that omits the foreign account reporting obligation, gets the primary tax number wrong by over $400, and fails to read the portfolio holdings the user asked about.


Response 2 — FAIL (320 / 795)

Response 2 is the weakest of the three. It compounds the standard deduction error with a second error on the Child Tax Credit, gives only a cursory nod to foreign account reporting, and offers the most generic portfolio advice of any response.

Critical failures

  • FBAR/FATCA mention is too weak to earn credit. The response states the foreign accounts "likely require foreign account reporting" but never names FBAR (FinCEN Form 114) or Form 8938, never states the $10,000 FBAR threshold or the $100,000 Form 8938 threshold, and never mentions penalties. Telling a user that something "likely" requires reporting without naming what or why is not actionable — it is a brush-off, not a finding. It loses Positive #1 (+100).
  • Uses the wrong standard deduction AND the wrong Child Tax Credit. The response applies $30,000 instead of $31,500 for the standard deduction, and $4,400 instead of $4,000 for the CTC — it trusted the draft on both counts. The combined errors produce a balance due of $1,114.49 instead of $1,941.99 — off by $827.50, the largest error of any response. A user acting on this would underpay by over $800. It loses Positive #4 (+55) and Positive #8 (+45), and triggers Negative #33 (−60) and Negative #32 (−80).
  • Portfolio advice is entirely generic. The response recommends "raise your GM 401(k) contribution to the annual max" from $15,300 to $23,500. This advice requires zero file analysis — it is the same recommendation that would be given to any person with a 401(k). The response never examines the 401(k) holdings, never mentions the GM stock concentration, never looks at the IRA or foreign accounts, and uses a generic 7% return assumption instead of the portfolio's documented returns. It loses Positive #3 (+70).

Strengths

  • Correctly includes all $8,480 in foreign interest and identifies the $656 Foreign Tax Credit.
  • Correctly rejects the family support and union dues deductions, and catches the draft's fictional "other dependents" credit.
  • Handles the classify-before-count trap correctly in the cash-flow section.
  • Cites source documents and fields for most figures.

Bottom line: Response 2 fails because it compounds two tax constant errors — wrong standard deduction and wrong CTC, both copied from the draft — producing a balance due off by over $800, gives a legally inadequate mention of foreign account reporting that names no forms or thresholds, and offers portfolio advice indistinguishable from a generic internet search. A careful reviewer would not ship an answer that gets the primary tax number wrong by $827.50 and cannot name the forms Amanuel needs to file.

Ground TruthVerified reference calculations+

Persona 43 -- Amanuel T. Ghebremariam -- Task 5 Ground Truth

Version: 2.0 — Corrected 2026-07-17

Changes from v1: Fixed standard deduction ($30,000→$31,500 per IRS OBBBA + project methodology), cascading tax/bottom-line recalc, exact payroll figures used throughout. Tax brackets and SALT cap verified against IRS publications, not the draft tax tracker.

Source Documents

All source files are in the private evaluation source set. Every figure below is traced to a specific source document and field. The canonical filenames from the prompt are used throughout.

Tax Constants (Verified Against Authoritative Sources)

Tax constants used in all calculations below. None are sourced from the user's draft tax tracker:

Constant Value Source
2025 MFJ standard deduction $31,500.00 Project methodology + IRS (OBBBA July 2025): "2025 standard deduction is $15,750 single / $31,500 MFJ"
2025 SALT cap (MFJ) $40,000.00 OBBBA (July 2025), raised from TCJA $10,000. Phases down by 30% of MAGI above $500k, floored at $10k at $600k MAGI. Amanuel's MAGI ($146,500) is well under the phaseout threshold.
2025 MFJ tax brackets 10%/12%/22%/24%/32%/35%/37% IRS Rev. Proc. 2024-40 (verified against IRS.gov)
2025 Child Tax Credit $2,000/child under 17 Per project standard
2025 IRA deduction phase-out (active participant, MFJ) $126,000–$146,000 MAGI IRS 2025

CATEGORY 1: TAX PLANNING

1.1 Income Computation

W-2 Wages (Box 1)

Source: W2_GeneralMotors_2025.pdf, Box 1

Value: $135,900.00

Reconciliation of gross ($153,000) to Box 1 ($135,900):

  • Gross salary: $153,000.00 (payroll_earnings_statements_2025.pdf, Annual Summary, line "Gross earnings (regular salary)")
  • Pre-tax 401(k) Code D: -$15,300.00 (W-2 Box 12a, code D)
  • Pre-tax medical Section 125: -$1,800.00 (payroll_earnings_statements_2025.pdf, Annual Summary, line "Pre-tax medical (Section 125)")
  • Box 1 = $153,000 - $15,300 - $1,800 = $135,900.00 — VERIFIED
W-2 Key Fields (All Verified)

Source: W2_GeneralMotors_2025.pdf

Box Description Value
1 Wages, tips, other comp. $135,900.00
2 Federal income tax withheld $9,186.00
3 Social security wages $151,200.00
4 Social security tax withheld $9,374.40
5 Medicare wages and tips $151,200.00
6 Medicare tax withheld $2,192.40
12a Code D (401k) $15,300.00
12b Code DD (employer health) $21,640.00
13 Retirement plan Checked [X]
14 Other — UAW DUES $2,200.00
16 State wages (MI) $135,900.00
17 State income tax (MI) $4,790.00
18 Local wages $135,900.00
19 Local income tax (Detroit) $3,204.00
20 Locality name DETROIT

Note: W-2 Box 17 rounds MI withholding to $4,790.00. The payroll statement shows the exact figure of $4,789.68. For SALT calculation we use the exact payroll figure ($4,789.68); for tax-return purposes the W-2 figure ($4,790.00) is what was reported.

Payroll Reconciliation

Source: payroll_earnings_statements_2025.pdf, Annual Summary

Line Value
Gross earnings (regular salary) $153,000.00
Pre-tax 401(k) — Code D $15,300.00
Pre-tax medical (Section 125) $1,800.00
Federal income tax withheld $9,186.00
Social Security tax withheld $9,374.40
Medicare tax withheld $2,192.40
Michigan income tax withheld $4,789.68
Detroit city tax withheld $3,204.00
UAW union dues (post-tax) $2,200.08
Net pay deposited $104,953.44

Reconciliation checks:

  • SS wages: $153,000 - $1,800 = $151,200.00 ✓ (matches W-2 Box 3/5)
  • Net pay: $153,000 - $15,300 - $1,800 - $9,186 - $9,374.40 - $2,192.40 - $4,789.68 - $3,204.00 - $2,200.08 = $104,953.44 ✓
  • 24 deposits × $4,373.06 = $104,953.44 ✓
Interest Income

Source: bank_checking_statements_2025.pdf, page 1 (Annual Account Statement)

  • Comerica checking interest: $519.96 (12 monthly credits × $43.33)
  • Verified: 12 interest payment lines in transactions_checking_2025.csv, each $43.33
  • The draft return shows $520.00 — minor rounding, $0.04 difference

Source: foreign_cbe_ethiopia_statement_2025.pdf, "Certificate of Interest Paid" page

  • CBE Ethiopia gross interest: $6,180.00 (ETB 971,805.00 @ 157.25 ETB/USD)
  • Tax withheld at source: ETB 97,180.50 (10%) = $618.00
  • Net interest credited: ETB 874,624.50
  • This is NOT included in the draft return's interest line

Source: foreign_cber_eritrea_statement_2025.pdf, "Certificate of Interest Paid" page

  • CBER Eritrea gross interest: $760.00 (ERN 11,400.00 @ 15.00 ERN/USD)
  • Tax withheld at source: ERN 570.00 (5%) = $38.00
  • Net interest credited: ERN 10,830.00
  • This is NOT included in the draft return's interest line

Source: foreign_emiratesnbd_uae_statement_2025.pdf, year-end summary

  • Emirates NBD UAE gross interest: $1,540.00 (AED 5,655.65; USD equivalent stated on statement)
  • Tax withheld at source: $0.00 (UAE does not withhold)
  • This is NOT included in the draft return's interest line

Total interest income (all sources): $519.96 + $6,180.00 + $760.00 + $1,540.00 = $8,999.96

Dividend Income

Source: brokerage_taxable_1099_2025.pdf, Year-End Investment Report

  • Total ordinary dividends: $1,600.00
  • Qualified dividends: $1,450.00
  • Non-qualified (SWVXX money market): $150.00
  • No securities sold in 2025 — no 1099-B issued
Total Income (Correct)
Source Amount Verified Against
W-2 wages (Box 1) $135,900.00 W-2 Box 1, payroll reconciliation
Comerica interest $519.96 12 × $43.33 in checking CSV
CBE Ethiopia interest $6,180.00 CBE certificate of interest
CBER Eritrea interest $760.00 CBER certificate of interest
Emirates NBD UAE interest $1,540.00 ENBD year-end summary
Schwab dividends $1,600.00 Brokerage 1099-DIV
Total income $146,499.96
Draft Return Error on Income

The draft return (tax_tracker_DRAFT_2025.xlsx, "Income (DRAFT)" sheet) shows:

  • Wages: $135,900.00 — CORRECT
  • Interest — Comerica: $520.00 — APPROXIMATELY CORRECT ($519.96 actual)
  • Dividends — Schwab: $1,600.00 — CORRECT
  • Interest — overseas accounts: BLANK — ERROR: $8,480.00 missing
  • Total income: $138,020.00 — ERROR: should be $146,499.96

The draft return omits $8,480.00 in foreign interest income ($6,180 + $760 + $1,540).


1.2 Adjusted Gross Income (AGI)

No adjustments to income are present in the draft return. The only potential adjustment would be:

  • IRA deduction: Amanuel is an active participant in an employer plan (W-2 Box 13 checked). For MFJ 2025, the traditional IRA deduction phase-out range is $126,000–$146,000 MAGI. His MAGI of $146,499.96 exceeds $146,000, so no traditional IRA deduction is available.
  • Student loan interest: Not present in documents.

Correct AGI = $146,499.96


1.3 Deductions — Itemized vs. Standard

2025 Standard Deduction (MFJ)

Source: Project methodology (methodology_pages_35_44.txt, "Source Beats Memory" trap #5)

Value: $31,500.00

"2025 standard deduction is $15,750 single / $31,500 MFJ — models keep typing $15,000 / $30,000."

This is a deliberate Source Beats Memory trap. Models trained on prior-year data will default to $30,000. The correct project-defined figure is $31,500.

Draft Return Itemized Deductions

Source: tax_tracker_DRAFT_2025.xlsx, "Deductions (DRAFT)" sheet

Item Draft Amount Correct?
Mortgage interest $8,213.89 YES — matches 1098 Box 1
State + local + property taxes $11,794.00 YES — under $40k SALT cap, fully deductible
Money sent to family (Eritrea) $24,000.00 NO — NOT deductible
Union dues (UAW) $2,200.00 NO — NOT deductible
Total itemized (draft) $46,207.89 OVERSTATED by $26,200+
Correct Itemized Deductions

Mortgage Interest: $8,213.89

  • Source: mortgage_statement_1098_2025.pdf, Form 1098 Box 1
  • Verified: Total interest paid in 2025 = $8,213.89

SALT (State and Local Taxes) — Fully Deductible Under $40,000 Cap:

Component Amount Source
Michigan income tax withheld $4,789.68 payroll_earnings_statements_2025.pdf, Annual Summary
Detroit city tax withheld $3,204.00 W-2 Box 19 / payroll Annual Summary
Property taxes (Wayne County + Detroit) $3,800.00 mortgage_statement_1098_2025.pdf, Form 1098 Box 10
Total SALT paid $11,793.68
2025 SALT cap (MFJ) $40,000.00 OBBBA (July 2025); phases down by 30% of MAGI above $500k, floored at $10k at $600k
Deductible SALT $11,793.68 Fully deductible (under $40k cap)

Family Support ($24,000): $0.00 — NOT deductible

  • Personal gifts/transfers to family members are not deductible. These are not charitable contributions (no 501(c)(3) recipient) and not a business expense.

Union Dues ($2,200): $0.00 — NOT deductible

  • UAW dues are post-tax (payroll_earnings_statements_2025.pdf, Annual Summary: "UAW union dues (post-tax)"). Miscellaneous itemized deductions subject to 2% AGI floor were suspended by TCJA through 2025.

Total Correct Itemized Deductions: $8,213.89 + $11,793.68 = $20,007.57

Standard Deduction Comparison
  • Standard deduction MFJ 2025: $31,500.00
  • Correct itemized: $20,007.57
  • Standard deduction is larger by $11,492.43

Conclusion: Amanuel should take the standard deduction of $31,500.00, not itemize.


1.4 Taxable Income

Line Correct Value Draft Value
Total income $146,499.96 $138,020.00
Standard deduction ($31,500.00) N/A (itemized $46,207.89)
Taxable income $114,999.96 $91,812.11

1.5 Tax Before Credits — 2025 MFJ Tax Brackets

2025 MFJ tax brackets (IRS Rev. Proc. 2024-40):

  • 10%: $0 to $23,850
  • 12%: $23,851 to $96,950
  • 22%: $96,951 to $206,700
  • 24%: $206,701 to $394,600
  • 32%: $394,601 to $501,050
  • 35%: $501,051 to $751,600
  • 37%: Over $751,600
Correct Tax Calculation

Taxable income: $114,999.96

Bracket Rate Amount in Bracket Tax
$0 – $23,850 10% $23,850.00 $2,385.00
$23,851 – $96,950 12% $73,100.00 $8,772.00
$96,951 – $114,999.96 22% $18,049.96 $3,970.99
Total $114,999.96 $15,127.99
Draft Return Tax Calculation (Verified)

The draft shows tax before credits of $10,366.45 on taxable income of $91,812.11.

Verification of what the draft's tax SHOULD be on its own numbers:

  • 10% on $23,850 = $2,385.00
  • 12% on ($91,812.11 - $23,850) = 12% on $67,962.11 = $8,155.45
  • Total = $2,385.00 + $8,155.45 = $10,540.45

The draft's $10,366.45 does not match the correct calculation of $10,540.45 even on its own numbers. The draft's tax before credits is internally inconsistent — it's off by $174.00.


1.6 Child Tax Credit

Source: Persona brief states 2 children (Danait, Nahom), both US-born

2025 CTC: $2,000 per qualifying child under 17

Correct CTC: 2 × $2,000 = $4,000.00

Draft shows:

  • Child Tax Credit: $4,400.00 (count: 2) — ERROR: $400 over ($2,200/child is wrong; correct is $2,000/child)
  • Credit for other dependents: $1,000.00 (count: 2) — ERROR: fictional (there are only 2 children, both qualify for full CTC; no other dependents exist)

Correct total credits: $4,000.00 (not $5,400.00 as draft shows)


1.7 Tax After Credits

Correct Calculation
Line Amount
Tax before credits $15,127.99
Child Tax Credit (2 × $2,000) ($4,000.00)
Tax after credits $11,127.99
Draft Return Calculation
Line Amount
Tax before credits $10,366.45
Credits (CTC $4,400 + Other $1,000) ($5,400.00)
Tax after credits $4,966.45

1.8 Refund or Balance Due

Correct
Line Amount
Tax after credits $11,127.99
Federal tax withheld (W-2 Box 2) ($9,186.00)
Balance due $1,941.99
Draft Return
Line Amount
Tax after credits $4,966.45
Federal tax withheld ($9,186.00)
Expected refund $4,219.55

The draft return is WRONG by $6,161.54. Amanuel actually owes $1,941.99, not receiving a refund of $4,219.55. The swing = $4,219.55 + $1,941.99 = $6,161.54.


1.9 State and Local Tax

Michigan State Tax
  • Rate: 4.25% flat
  • MI taxable wages: $135,900.00 (W-2 Box 16)
  • MI tax: $135,900 × 4.25% = $5,775.75
  • MI withheld: $4,789.68 (payroll Annual Summary; W-2 Box 17 rounds to $4,790.00)
  • Balance due to Michigan: $986.07
Detroit City Tax
  • Rate: 2.4% for residents
  • Detroit taxable wages: $135,900.00 (W-2 Box 18)
  • Detroit tax: $135,900 × 2.4% = $3,261.60
  • Detroit withheld: $3,204.00 (W-2 Box 19)
  • Balance due to Detroit: $57.60
Combined State/Local
  • Total state + local balance due: $1,043.67

1.10 Foreign Account Reporting

FBAR (FinCEN Form 114)

Filing threshold: Aggregate foreign accounts > $10,000 at any time during year

Amanuel's foreign accounts (12/31/2025 balances):

Account Balance (USD) Source
CBE Ethiopia $58,000.00 foreign_cbe_ethiopia_statement_2025.pdf, Account Summary
Emirates NBD UAE $42,000.00 foreign_emiratesnbd_uae_statement_2025.pdf, Account Summary
CBER Eritrea $21,200.00 foreign_cber_eritrea_statement_2025.pdf, Account Summary
Total $121,200.00

FBAR filing is REQUIRED — total far exceeds $10,000 threshold. Amanuel has never filed.

FATCA Form 8938

Filing threshold for MFJ living in US: $100,000 aggregate foreign financial assets on last day of year, or $150,000 at any time during year.

Amanuel's foreign financial assets total $121,200.00 on 12/31/2025.

Form 8938 filing is REQUIRED for MFJ ($100,000 year-end threshold exceeded).

Penalty Risk
  • Non-willful FBAR violation: up to $10,000 per violation (adjusted for inflation)
  • Non-willful Form 8938 violation: up to $10,000 per violation
  • Amanuel should consider Streamlined Filing Compliance Procedures for non-willful offshore non-compliance
Foreign Tax Credit

Foreign taxes withheld on interest (verified from each bank's certificate of interest):

Country Gross Interest Withholding Rate Tax Withheld
Ethiopia $6,180.00 10% $618.00
Eritrea $760.00 5% $38.00
UAE $1,540.00 0% $0.00
Total $8,480.00 $656.00

Amanuel may be eligible for a Foreign Tax Credit on Form 1116 for the $656.00 in foreign taxes withheld. This would reduce his US tax liability dollar-for-dollar. The draft return does not address this at all.


1.11 Inheritance — Tax Implications

Source: property_inheritance_asmara_2025.pdf

  • 1/3 undivided interest in family home in Asmara, Eritrea
  • Fair value: $30,000 USD (ERN 450,000 @ 15.00 ERN/USD)
  • Occupied by decedent's widow, not income-producing
  • Decedent: Tesfamariam Ghebremariam, died 11 August 2023
  • Intestate succession; devolved equally to three surviving sons

US tax treatment: Inheritances are generally not taxable income to the recipient. No reporting required on Form 1040. The property is real estate (not a financial account), so no FBAR/FATCA reporting is needed for this asset. No step-up in basis applies to foreign real estate the same way it does for US securities, but since this is an inheritance (not a sale), there is no current US tax event.


1.12 Summary of Draft Return Errors

Issue Draft Correct Impact
Foreign interest income $0 $8,480.00 Understated income
Total income $138,020.00 $146,499.96 Understated by $8,479.96
Family support deduction $24,000.00 $0.00 Not deductible
Union dues deduction $2,200.00 $0.00 Not deductible (TCJA)
SALT deduction $11,794.00 $11,793.68 Correct (under $40k cap)
Itemized deductions total $46,207.89 $20,007.57 Overstated by $26,200.32
Deduction choice Itemized Standard ($31,500) Wrong choice
Taxable income $91,812.11 $114,999.96 Understated by $23,187.85
Tax before credits $10,366.45 $15,127.99 Understated; also internally inconsistent
Child Tax Credit $4,400.00 $4,000.00 Overstated by $400
Other dependents credit $1,000.00 $0.00 Fictional
Tax after credits $4,966.45 $11,127.99 Understated
Refund/balance due +$4,219.55 (refund) -$1,941.99 (owed) $6,161.54 error
FBAR filing Not addressed Required Penalty risk
Form 8938 filing Not addressed Required Penalty risk
Foreign tax credit Not addressed $656 available Missed tax savings
MI state tax Not addressed $986.07 due Underwithheld
Detroit city tax Not addressed $57.60 due Underwithheld

1.13 Common Model Failure Modes — Tax

  1. Using draft return numbers without verification — Models that accept the draft's $138,020 income without adding foreign interest will be wrong.
  2. Treating family support as deductible — The $24,000 wire to Eritrea is a common trap. It is a personal gift/transfer, not deductible.
  3. Treating union dues as deductible — UAW dues are post-tax and not deductible in 2025 (TCJA suspended miscellaneous itemized deductions).
  4. Source Beats Memory: Standard deduction — Models default to $30,000 (memorized); the project-defined 2025 MFJ figure is $31,500. Using $30,000 makes taxable income and tax wrong.
  5. SALT cap — Models must apply the correct $40,000 SALT cap (per project source documents). The full $11,793.68 in SALT paid is under the cap and fully deductible. Models using the stale $10,000 TCJA cap will incorrectly limit the deduction.
  6. Standard vs. itemized comparison — The correct choice is standard deduction ($31,500), not itemized ($20,007.57).
  7. CTC amount — $2,000/child, not $2,200 as the draft implies.
  8. Foreign income inclusion — Foreign interest is taxable to US residents on worldwide income.
  9. FBAR/FATCA — Many models miss the reporting requirements entirely or confuse the thresholds.
  10. Foreign tax credit — Available ($656) but not claimed in draft.
  11. State/local tax underpayment — MI and Detroit taxes are underwithheld (~$1,044 combined).

CATEGORY 2: SUBSCRIPTIONS / CASH-FLOW

2.1 Recurring Subscriptions (from Checking CSV)

Identified Monthly Subscriptions

Source: transactions_checking_2025.csv (540 transactions, 12 months)

Subscription Monthly Cost Annual Cost Verification
T-Mobile PCS SVC $165.00 $1,980.00 12 occurrences, every month
Netflix $15.49 $185.88 12 occurrences, every month
Spotify USA $11.99 $143.88 12 occurrences, every month
NYTimes Digital $17.00 $204.00 12 occurrences, every month
Apple iCloud $2.99 $35.88 12 occurrences, every month
Planet Fitness $49.00 $588.00 12 occurrences, every month
Comcast Xfinity $89.00 $1,068.00 12 occurrences, every month
DTE Energy ~$220/mo avg ~$2,640.00 12 occurrences, seasonal variation
City of Detroit Water ~$105/mo avg ~$1,260.00 12 occurrences, varies by month
Total subscriptions/utilities ~$675/mo ~$8,105/yr
Other Recurring Monthly Items
Item Monthly Annual Verification
Mortgage (P&I + escrow) $1,419.46 $17,033.52 12 occurrences, 3rd of each month
Schwab brokerage ACH $500.00 $6,000.00 12 occurrences, ~15th of each month
Family wire to Eritrea $2,000.00 $24,000.00 12 occurrences, ~12th of each month
Chase credit card autopay ~$1,347/mo avg $16,170.61 12 occurrences, ~22nd of each month

2.2 Cash Flow Analysis

Annual Income (Net Deposits)

Source: bank_checking_statements_2025.pdf, page 1 (Annual Account Statement)

Component Amount Source
Total deposits & credits $105,616.27 Bank statement, page 1
Payroll deposits $104,953.44 24 × $4,373.06 (verified in CSV)
Interest credits $519.96 12 × $43.33 (verified in CSV)
Meijer reversal $142.87 One-time, CSV line 260

Reconciliation: $104,953.44 + $519.96 + $142.87 = $105,616.27 ✓

Annual Expenses (From Checking)

Source: bank_checking_statements_2025.pdf, page 1

  • Total withdrawals & debits: $97,979.73
Net Cash Flow from Checking

$105,616.27 - $97,979.73 = $7,636.54 positive cash flow

Ending Balance Verification
  • Beginning balance (Jan 1, 2025): $9,240.00
  • Net change: +$7,636.54
  • Ending balance (Dec 31, 2025): $16,876.54 — VERIFIED (matches bank statement and CSV final line)
Linked Savings

Source: bank_checking_statements_2025.pdf, page 1

  • Comerica Statement Savings ****5120: $3,180.00
Total Liquid Assets

$16,876.54 + $3,180.00 = $20,056.54


2.3 Chase Credit Card Spending by Category (2025)

Source: credit_card_statements_2025.pdf, year-to-date summary + monthly statements

Category Annual Monthly Avg
Groceries $3,061.81 $255.15
Dining $2,869.07 $239.09
Fuel $1,852.38 $154.37
Pharmacy $1,757.41 $146.45
Online Retail $1,236.40 $103.03
Electronics $1,112.03 $92.67
Retail $1,006.88 $83.91
Travel $1,002.06 $83.51
Warehouse (Costco) $926.53 $77.21
Home Improvement $659.66 $54.97
Education (Udemy) $642.35 $53.53
Total $16,126.58 $1,343.88

Card summary verification:

  • Total purchases: $16,126.58 ✓
  • Total payments/credits: $16,170.61 ✓
  • Total interest charged: $0.00 ✓
  • Total fees charged: $0.00 ✓
  • Statement balance 12/31/2025: $1,136.39 ✓
  • Credit limit: $22,000.00
  • Card paid in full monthly by autopay — no interest charges

2.4 Combined Monthly Spending (Checking + Credit Card)

Critical classification note: The Chase autopay from checking (~$1,347/mo) is a TRANSFER that pays off credit card purchases. The actual spending is on the Chase card, not the autopay. Adding both would double-count. The combined analysis below uses Chase spending categories (not the autopay) plus non-Chase checking expenses.

Category Checking (Monthly) Credit Card (Monthly) Total Monthly
Mortgage $1,419 $0 $1,419
Family wire $2,000 $0 $2,000
Schwab investment $500 $0 $500
Utilities (DTE, Comcast, T-Mobile, Water) $579 $0 $579
Subscriptions (Netflix, Spotify, NYT, iCloud, Planet Fitness) $96 $0 $96
Groceries ~$433 $255 $688
Dining ~$233 $239 $472
Fuel ~$133 $154 $287
Pharmacy ~$50 $146 $196
Retail/Online ~$292 $240 $532
Travel $0 $84 $84
Education $0 $54 $54
Electronics $0 $93 $93
Home Improvement $0 $55 $55
Coffee shops (Biggby, Starbucks, Dunkin, Tim Hortons) ~$150 $0 $150
ATM/Cash ~$97 $0 $97
Total ~$5,982 ~$1,320 ~$7,302

Note: Checking categories are approximate (derived from merchant names in CSV). Credit card categories are from Chase's own categorization.


2.5 True Surplus/Deficit

  • Net monthly take-home pay: $104,953.44 / 12 = $8,746.12
  • Total monthly outflow (all categories including investments & family support): ~$7,302
  • Monthly surplus: ~$1,444
  • Annual surplus: ~$17,328

This surplus accumulates in checking (ending $16,876.54 vs beginning $9,240.00 = +$7,636.54) and the linked savings ($3,180.00). The difference between the annual surplus (~$17,328) and the checking accumulation ($7,636.54) is partly explained by the savings account growth and timing/categorization differences.


2.6 Key Cash-Flow Observations

  1. The $2,000/month family wire is the single largest expense at ~27% of total monthly outflow. This is a non-negotiable cultural/family obligation.
  2. The $500/month Schwab ACH is forced savings/investment — it goes to the taxable brokerage and is not consumption.
  3. Chase card is paid in full monthly — no interest, good discipline. Average monthly spend ~$1,344.
  4. Dining out is significant — combined checking + credit card dining is ~$472/month ($5,664/year).
  5. Coffee shops (Biggby, Starbucks, Dunkin, Tim Hortons) appear frequently — estimated $150-200/month from checking.
  6. The linked savings account ($3,180) is small relative to total cash flow. Emergency fund (~$20,057 total liquid) covers about 2.7 months of expenses.
  7. No credit card debt — utilization averages 5-8%, well below the 30% threshold.

2.7 Common Model Failure Modes — Cash Flow

  1. Double-counting the Chase autopay — The checking account shows the payment (~$1,347/mo), and the Chase statement shows the purchases. Models that add both will double-count ~$16,170/year.
  2. Treating the Schwab ACH as an expense — It's an investment transfer, not consumption. Should be classified separately.
  3. Missing the family wire — $24,000/year is a major cash outflow that must be included in any cash flow analysis.
  4. Ignoring the linked savings account — The $3,180 in savings is part of total liquidity.
  5. Not reconciling net pay with deposits — The $104,953.44 net pay should match the payroll deposits in checking.

CATEGORY 3: PORTFOLIO ANALYSIS

3.1 Retirement Accounts

GM 401(k) (Fidelity)

Source: retirement_401k_statement_2025.pdf

Item Amount Verified
Beginning balance (01/01/2025) $160,200.00 Statement "Beginning vested balance"
Employee pre-tax contributions $15,300.00 Statement "Your pre-tax contributions (Code D)"
Employer match $7,650.00 Statement "Employer matching contributions" (100% of first 5%)
Change in market value $22,150.00 Statement "Change in market value (earnings)"
Ending balance (12/31/2025) $205,300.00 Statement "Ending vested balance"
Personal rate of return +13.1% Statement

Contribution rate: 10% pre-tax. Employer match: 100% of first 5%. $153,000 × 5% = $7,650 ✓

Holdings as of 12/31/2025:

Fund Symbol Shares Price Value % Verified
Fidelity 500 Index FXAIX 428.239 $201.35 $86,226.00 42.0% 428.239 × $201.35 = $86,225.92 ≈ $86,226 ✓
Fidelity Freedom 2045 FFFGX 3,343.648 $18.42 $61,590.00 30.0% 3,343.648 × $18.42 = $61,589.99 ≈ $61,590 ✓
Fidelity US Bond Index FXNAX 2,708.954 $10.61 $28,742.00 14.0% 2,708.954 × $10.61 = $28,742.00 ✓
Vanguard Extended Market VIEIX 100.902 $182.90 $18,455.00 9.0% 100.902 × $182.90 = $18,454.98 ≈ $18,455 ✓
GM Common Stock GMSTK 197.447 $52.10 $10,287.00 5.0% 197.447 × $52.10 = $10,286.99 ≈ $10,287 ✓
Total $205,300.00 100%
Rollover IRA (Fidelity)

Source: ira_rollover_statement_2025.pdf

Item Amount Verified
Beginning balance (01/01/2025) $48,520.00 Statement
Contributions $0.00 Statement
Change in market value $6,180.00 Statement
Ending balance (12/31/2025) $54,700.00 Statement

No contributions or distributions in 2025. Under age 73 — no RMD applies.

Holdings as of 12/31/2025:

Fund Symbol Shares Price Value % Verified
Fidelity 500 Index FXAIX 114.726 $201.35 $23,100.00 42.2% 114.726 × $201.35 = $23,100.08 ≈ $23,100 ✓
Fidelity Total International FTIHX 740.214 $14.05 $10,400.00 19.0% 740.214 × $14.05 = $10,400.01 ≈ $10,400 ✓
Fidelity US Bond Index FXNAX 1,187.559 $10.61 $12,600.00 23.0% 1,187.559 × $10.61 = $12,600.00 ✓
Fidelity Govt Money Market SPAXX 8,600.000 $1.00 $8,600.00 15.7% 8,600 × $1.00 = $8,600 ✓
Total $54,700.00 100%
Total Retirement Assets
Account Value
GM 401(k) $205,300.00
Rollover IRA $54,700.00
Total retirement $260,000.00

3.2 Taxable Brokerage (Schwab)

Source: brokerage_taxable_1099_2025.pdf

Item Amount Verified
Beginning value (01/01/2025) $52,180.00 Statement
Cash transfers in ($500/mo) $6,000.00 Statement (12 × $500 from Comerica checking)
Dividends received $1,600.00 Statement
Change in market value -$1,207.35 Statement
Ending value (12/31/2025) $58,572.65 Statement
Unrealized gain $17,681.65 Statement

No securities sold in 2025 — no 1099-B issued.

Holdings as of 12/31/2025:

Holding Symbol Shares Price Value % Verified
Vanguard Total Stock Market ETF VTI 118.000 $291.40 $34,385.20 58.7% 118 × $291.40 = $34,385.20 ✓
Schwab US Dividend Equity ETF SCHD 205.000 $29.85 $6,119.25 10.4% 205 × $29.85 = $6,119.25 ✓
Apple Inc. AAPL 42.000 $236.20 $9,920.40 16.9% 42 × $236.20 = $9,920.40 ✓
Microsoft Corp. MSFT 18.000 $452.60 $8,146.80 13.9% 18 × $452.60 = $8,146.80 ✓
Schwab Money Market SWVXX 1.000 $1.00 $1.00 0.0% 1 × $1.00 = $1.00 ✓
Total $58,572.65 100%

Dividend breakdown: $1,600.00 total ($1,450 qualified + $150 non-qualified from SWVXX)


3.3 Foreign Accounts

Account Country Currency Balance (local) Balance (USD) Source
CBE Ethiopia Ethiopia ETB 9,120,500.00 $58,000.00 foreign_cbe_ethiopia_statement_2025.pdf
Emirates NBD UAE AED 154,245.00 $42,000.00 foreign_emiratesnbd_uae_statement_2025.pdf
CBER Eritrea Eritrea ERN 318,000.00 $21,200.00 foreign_cber_eritrea_statement_2025.pdf
Total foreign $121,200.00

Foreign interest earned in 2025 (all verified from certificates of interest):

Country Gross Interest (USD) Tax Withheld Net
Ethiopia $6,180.00 $618.00 (10%) $5,562.00
Eritrea $760.00 $38.00 (5%) $722.00
UAE $1,540.00 $0.00 (0%) $1,540.00
Total $8,480.00 $656.00 $7,824.00

3.4 Home Equity

Source: mortgage_statement_1098_2025.pdf

Item Amount Verified
Original loan amount $158,000.00 Statement
Principal balance 01/01/2025 $132,990.00 1098 Box 2
Principal balance 12/31/2025 $129,530.41 Statement ($132,990 - $3,459.59 principal reduction)
Total principal reduction in 2025 $3,459.59 Statement
Interest rate 6.250% fixed Statement
Monthly P&I $972.79 Statement
Monthly escrow $446.67 Statement
Total monthly payment $1,419.46 Statement ($972.79 + $446.67)
Estimated property value $190,000.00 Statement
Home equity $60,469.59 $190,000 - $129,530.41
Loan-to-value 68.2% $129,530.41 / $190,000

3.5 Inheritance

Source: property_inheritance_asmara_2025.pdf

  • 1/3 undivided interest in Asmara family home
  • Fair value: $30,000.00 USD (ERN 450,000 @ 15.00 ERN/USD)
  • Full property value: $90,000.00 USD (ERN 1,350,000)
  • Not income-producing, occupied by decedent's widow
  • Not a financial account for FBAR/FATCA purposes
  • Three equal heirs: Amanuel, Yosief, Tesfay

3.6 Other Assets

Asset Value Source
Comerica checking $16,876.54 bank_checking_statements_2025.pdf
Comerica linked savings $3,180.00 bank_checking_statements_2025.pdf
Total cash/liquid $20,056.54

3.7 Total Net Worth

Category Value
Assets
GM 401(k) $205,300.00
Rollover IRA $54,700.00
Schwab taxable brokerage $58,572.65
Comerica checking $16,876.54
Comerica savings $3,180.00
Home equity ($190,000 - $129,530.41) $60,469.59
CBE Ethiopia $58,000.00
Emirates NBD UAE $42,000.00
CBER Eritrea $21,200.00
Inheritance (Asmara) $30,000.00
Total assets $550,298.78
Liabilities
Mortgage balance ($129,530.41)
Chase credit card (statement balance 12/31/2025) ($1,136.39)
Total liabilities ($130,666.80)
Net worth $419,631.98

3.8 Asset Allocation Analysis

Combined Retirement Allocation (401k + IRA = $260,000)
Asset Class 401(k) IRA Combined % of Retirement
US Large Cap (FXAIX) $86,226 $23,100 $109,326 42.0%
Target Date 2045 (FFFGX) $61,590 $0 $61,590 23.7%
US Bonds (FXNAX) $28,742 $12,600 $41,342 15.9%
US Mid/Small (VIEIX) $18,455 $0 $18,455 7.1%
International (FTIHX) $0 $10,400 $10,400 4.0%
GM Stock (GMSTK) $10,287 $0 $10,287 4.0%
Cash/MM (SPAXX) $0 $8,600 $8,600 3.3%
Total $205,300 $54,700 $260,000 100%
Overall Portfolio (All Investable Assets)
Account Value % of Investable
401(k) $205,300.00 60.6%
Rollover IRA $54,700.00 16.2%
Taxable brokerage $58,572.65 17.3%
Checking/savings $20,056.54 5.9%
Total investable $338,629.19 100%
Overall Asset Allocation (Approximate)
Asset Class Estimated % Components
US Equities ~60% FXAIX, VTI, VIEIX, AAPL, MSFT, SCHD
Target Date / Balanced ~18% FFFGX (holds stocks + bonds internally)
Fixed Income / Bonds ~12% FXNAX
International Equity ~3% FTIHX
Cash / MM ~3% SPAXX, checking, savings
Single Stock (GM) ~3% GMSTK
Single Stocks (AAPL, MSFT) ~5% In taxable account

3.9 Portfolio Observations and Recommendations

Strengths
  1. Strong retirement savings — $260,000 at age 42 is on track. Contributing 10% + 5% employer match = 15% total.
  2. Low-cost index funds — FXAIX (0.015%), FXNAX (0.025%), VTI (0.03%) are excellent choices.
  3. No debt besides mortgage — Credit card paid in full monthly.
  4. Consistent investing — $500/month to taxable brokerage.
  5. Diversified across accounts — Retirement, taxable, foreign, real estate.
Concerns
  1. GM stock concentration — 5% of 401(k) in GMSTK + employment at GM = double risk. If GM struggles, both job and retirement are affected. Recommend reducing to 0-2%.
  2. Large cash position in IRA — $8,600 (15.7% of IRA) in SPAXX money market is conservative for a 42-year-old. Could be deployed into equities.
  3. Low international exposure — The 401(k) has no international equity fund. The IRA has FTIHX at 19%, but overall international is only ~3%. Target 15-20%.
  4. Single stock risk in taxable — AAPL (16.9%) and MSFT (13.9%) together are 30.8% of the taxable account. Consider diversifying.
  5. Foreign accounts are uninvested — $121,200 sitting in savings accounts abroad. While these serve family/diaspora purposes, the real returns may be eroded by currency fluctuations.
  6. Emergency fund — $20,057 in checking/savings is about 2.7 months of expenses. Recommended: 3-6 months ($21,000-$42,000).
Rebalancing Suggestions
  1. 401(k): Reduce GMSTK from 5% to 0-2%. Add Fidelity Total International Index (FTIHX) for 10-15% allocation.
  2. IRA: Reduce SPAXX from 15.7% to 5%. Deploy into FTIHX (increase to 25-30%) and FXAIX.
  3. Taxable: Direct all $500/month to VTI only. Consider selling AAPL/MSFT over time to reduce concentration (be mindful of capital gains taxes).

3.10 Common Model Failure Modes — Portfolio

  1. Forgetting the Rollover IRA — Models that only look at the 401(k) miss $54,700.
  2. Not combining holdings across accounts — FXAIX appears in both 401(k) and IRA; total US large cap exposure is understated if only one account is analyzed.
  3. Ignoring foreign accounts — $121,200 in foreign deposits is a significant part of net worth.
  4. Not flagging GM stock concentration — 5% of 401(k) in employer stock + employed by GM is a classic double-risk.
  5. Missing the dividend tax treatment — $1,450 of $1,600 dividends are qualified (lower tax rate).
  6. Not calculating net worth — Must include home equity, foreign accounts, and inheritance.
  7. Treating the inheritance as taxable income — It is not; it's an asset transfer.

HIDDEN HEADLINE FINDINGS

Finding 1: Amanuel Owes $1,942, Not Getting a $4,220 Refund

The draft return is catastrophically wrong. Amanuel believes he will receive a $4,219.55 refund. In reality, he owes $1,941.99. The $6,161.54 swing is driven by:

  • $8,480 in unreported foreign interest income
  • $28,000+ in improper deductions (family support $24,000, union dues $2,200)
  • Wrong choice of standard ($31,500) vs. itemized deduction
  • Overstated Child Tax Credit ($4,400 vs. $4,000) and fictional "other dependents" credit ($1,000)
  • Draft's own tax calculation is internally inconsistent (off by $174 even on its own numbers)

Finding 2: FBAR/FATCA Non-Compliance Risk

Amanuel has $121,200 in foreign accounts and has never filed FBAR (FinCEN Form 114) or Form 8938. Penalties for non-willful FBAR violations can reach $10,000 per violation. He should consider Streamlined Filing Compliance Procedures.

Finding 3: Foreign Tax Credit Worth $656

Amanuel paid $656 in foreign withholding taxes (Ethiopia $618, Eritrea $38). He can claim a Foreign Tax Credit on Form 1116 to reduce his US tax liability dollar-for-dollar. This is not reflected in the draft return.

Finding 4: State Tax Underpayment

Michigan tax is underwithheld by $986.07 and Detroit city tax by $57.60. Combined state/local balance due: ~$1,044.

Finding 5: $121,200 in Foreign Cash Earning Modest Returns

The foreign accounts earn 4.75-6% nominal interest, but after currency risk and inflation in those countries, real returns may be negative. Amanuel should evaluate whether these funds could be better deployed.

Finding 6: GM Stock Double Risk

Amanuel works for GM and holds 5% of his 401(k) in GM stock. If GM underperforms, both his salary and retirement savings are at risk. This is an uncompensated concentration risk.


TOLERANCE BANDS

Metric Ground Truth Acceptable Range Notes
Total income $146,499.96 $146,400–$146,600 Small rounding on interest
Standard deduction $31,500.00 $31,500 (fixed) 2025 MFJ per project methodology
SALT cap $40,000.00 $40,000 (fixed) 2025 MFJ per project source documents
Taxable income $114,999.96 $114,900–$115,100
Tax before credits $15,127.99 $15,050–$15,200 Bracket math
Child Tax Credit $4,000.00 $4,000 (fixed) 2 × $2,000
Tax after credits $11,127.99 $11,050–$11,200
Balance due (federal) $1,941.99 $1,870–$2,010
MI balance due $986.07 $980–$992
Detroit balance due $57.60 $55–$60
Foreign tax credit available $656.00 $650–$660
Net worth $419,631.98 $415,000–$425,000 Market value fluctuations
Monthly surplus ~$1,444 $1,300–$1,600 Varies by month
Foreign accounts total $121,200.00 $120,000–$122,000 Exchange rate dependent
Draft return error (swing) $6,161.54 $6,050–$6,270

This ground truth was compiled from the 14 source files in the private evaluation source set

Change Log from v1

Item v1 (Incorrect) v2 (Corrected) Reason
Standard deduction (MFJ 2025) $30,000.00 $31,500.00 Project methodology "Source Beats Memory" trap #5
MI state tax withheld $4,790.00 $4,789.68 Using exact payroll figure, not rounded W-2
SALT total paid $11,794.00 $11,793.68 $0.32 difference from exact payroll MI figure
Correct itemized deductions $20,007.89 $20,007.57 Minor: $0.32 from exact SALT figure
Taxable income $116,499.96 $114,999.96 Standard deduction +$1,500
Tax before credits $15,457.99 $15,127.99 $1,500 × 22% = $330 less tax
Tax after credits $11,457.99 $11,127.99 Cascade from above
Balance due (federal) $2,271.99 $1,941.99 Cascade from above
Draft return error (swing) $6,491.54 $6,161.54 Cascade from above
Task 05

Dual-Income Tax Reconciliation

Tax planning · Household employment

A dual-income household with household-employer obligations and a draft return that does not reconcile.

6 documents

The PromptWhat the simulated consumer asks, in their own voice+

Persona 30 — Niamh A. Okonkwo-Bauer, Task 6: Tax, Budgeting, and Subscriptions

Persona voice: 39-year-old DOJ trial attorney in DC, married to a hospital administrator, two kids in private school, full-time nanny. Smart, detail-oriented, frustrated that $420k feels like paycheck-to-paycheck. Did her own taxes, doesn't trust the result, wants to understand the numbers herself — not be handed off.

Assigned categories: Tax planning, Budgeting/spending, Subscriptions/cash-flow


The Prompt

I'm Niamh — 39, trial attorney at DOJ, married to Stefan who runs a department at Potomac General Health System. Between us we grossed $420,150 last year. We have two kids at Cathedral Heights, a full-time nanny named Marisol who's been with us for years and is absolutely not negotiable, a condo in DC, and we're still paying off my law school loans and Stefan's grad school loans through Cardinal. And I'm going to be honest with you: we make a fortune and there's nothing left on the 30th. I cannot tell you where $420,000 went.

I did our 2025 taxes myself using consumer tax software and it's showing a $1,333 federal refund. That number makes me nervous — it feels low for what we make, like maybe the software missed something, especially around Marisol's payroll. Stefan thinks I'm being paranoid and just wants to file it. Before we do, I need you to go through every form and schedule in our return — our W-2s, the Meridian 1099s, the FirstMeridian mortgage statement, the Cardinal 1098-Es, the household payroll records, everything — and tell me the exact corrected federal bottom line. Not a range, not "approximately." The exact dollar amount I should write on the check or the exact refund I should expect. And I want to see the math step by step with the source document and line reference for every major figure so I can follow it myself.

But the tax number is only half of what I need. The real question is whether we can actually afford to keep everything going as-is. Walk through our Tidewater checking account for the full year — all 340 transactions — and tell me what our real monthly cash flow looks like. Separate actual spending from the Capital One autopay and the transfers to and from savings and the Meridian brokerage, because I keep getting confused about whether a payment is new spending or just paying off last month's card balance. I want the true surplus or shortfall each month, and I want you to name the months where we're underwater and explain why. I know we have subscriptions running — Netflix, Spotify, Peloton, Disney+, NYTimes, iCloud — and I want every single one listed with what it costs annually, but I already know the real problem is bigger than streaming.

Here's what I need you to tell me: based on our actual cash flow and the corrected tax picture, can we afford to keep Marisol full-time, keep the kids at Cathedral Heights, keep the condo, and keep making the loan payments — all of it, as-is? And if we can't, what's the math that tells us where the real pressure is? Don't tell me to get rid of Marisol or pull the kids out of Cathedral Heights — those aren't on the table. Don't tell me to hire a CPA either. If I wanted a CPA I would have hired one. I want to understand this myself with the actual numbers from our actual files. Cite the file and the field for every dollar figure so I can check your work.

Critical ElementsThe findings a correct response must reach+

Persona 30 — Niamh A. Okonkwo-Bauer, Task 6: Critical Elements

Ground truth reference: persona_30_task6_ground_truth_v3.md

Rubric reference: persona_30_task6_rubric_v6.csv


  1. The corrected federal bottom line is a balance due of $8,288.24, not the $1,333.51 refund shown on the draft 1040 — a $9,621.75 swing against the taxpayer. The draft contains three errors: missing Schedule H ($7,753.20), missing Form 8959 ($1,421.19), and a wrongly claimed student loan interest deduction ($2,500 at 24% marginal = $600.00). The corrected total tax is $66,040.88 against total payments of $57,752.64. This element is critical because the entire affordability answer is wrong if the tax position is wrong.
  1. Schedule H household employment taxes of $7,753.20 are completely missing from the draft 1040. The nanny is paid $50,400.00 in cash wages on the books, triggering employer plus employee FICA of $7,711.20 (15.3% × $50,400) and FUTA of $42.00 (0.6% × $7,000). This is the single largest error in the draft, accounting for $7,753.20 of the $9,621.75 swing from refund to balance due.
  1. Form 8959 Additional Medicare Tax of $1,421.19 is completely missing from the draft 1040. Combined Medicare wages are $407,910.00 (Niamh $190,950.00 + Stefan $216,960.00), exceeding the $250,000 MFJ threshold by $157,910.00 at 0.9%. Stefan's employer already withheld $152.64 of this amount (0.9% on his $16,960.00 of wages over $200,000), which is a separate payment on line 25c, not part of the $57,600.00 Box 2 income tax withholding.
  1. The student loan interest deduction is $0.00, not the $2,500.00 claimed on the draft. At an AGI of $395,370.00, the deduction is fully phased out — the MFJ phaseout range ends at approximately $195,000. The actual interest paid is $10,202.00 ($7,474.00 Niamh at 5.05% + $2,728.00 Stefan at 4.40%), but none of it is deductible. Removing the $2,500.00 deduction increases AGI by $2,500.00 and line 16 tax by $600.00 at the 24% marginal rate.
  1. Total federal payments are $57,752.64, not the $57,600.00 shown on the draft. Box 2 income tax withholding is $57,600.00 ($33,049.92 Niamh + $24,550.08 Stefan) and goes on line 25a. The $152.64 of Additional Medicare Tax withholding from Stefan's Box 6 is a separate payment on line 25c. These are separate boxes for separate taxes — Box 2 is income tax, Box 6 is Medicare tax. Embedding the $152.64 in Box 2 understates payments and overstates the balance due.
  1. The corrected tax chain is: total income $395,370.00 → AGI $395,370.00 → itemized deductions $71,390.45 → taxable income $323,979.55 → line 16 tax $61,436.69 (ordinary income $301,619.55 taxed at 10%/12%/22%/24% = $58,082.69, plus QDI/LTCG $22,360.00 at 15% = $3,354.00) → credits $5,600.00 ($4,400.00 CTC + $1,200.00 CDCC) → subtotal $55,836.69 → other taxes $10,204.19 (Schedule H $7,753.20 + Form 8959 $1,421.19 + NIIT $1,029.80) → total tax $66,040.88.
  1. The SALT deduction of $36,369.48 is correct and fully allowed. The total comprises DC income tax withheld of $28,102.80 (Niamh $13,258.80 + Stefan $14,844.00 from W-2 Box 17) plus DC real estate taxes of $8,266.68 from the FirstMeridian Form 1098 escrow detail. Under current law the SALT cap is $40,000 for MFJ with phaseout beginning at $500,000 MAGI. Their MAGI of $395,370.00 is under the phaseout threshold. Applying the old $10,000 TCJA cap would incorrectly reduce itemized deductions by $26,369.48 and inflate the balance due by approximately $6,329.00.
  1. The Capital One credit card autopay of $30,630.71 across 12 monthly payments is a payment mechanism, not new spending. The checking account shows monthly autopay debits ranging from $2,132.39 to $3,042.67. These settle prior-month card balances and must not be added on top of the categorized purchases they pay off. Double-counting the autopay alongside underlying spending inflates stated outflows by approximately $30,000 per year.
  1. Combined student loan payments of $21,744.00 per year ($1,172.00/month Niamh at 5.05% + $640.00/month Stefan at 4.40%) are not visible anywhere in the 340 Tidewater checking account transactions. This is a material data gap representing 8.5% of net pay. The payments are being made from an untraced source — possibly a separate account, the credit card, or automatic debit from another bank.
  1. The household drained a net $15,748.46 from savings during the year to stay afloat. Savings transfers into checking totaled $27,000.00 ($12,000.00 in January, $15,000.00 in May) while only $11,251.54 was returned (a single December 31 transfer). The checking account still lost $2,950.00 over the year, falling from $21,450.00 to $18,500.00. Without savings infusions, the account would be near zero.
  1. The nanny fully-loaded employer cost is $54,540.60, comprising gross wages $50,400.00 + employer FICA $3,855.60 (6.2% SS + 1.45% Medicare on $50,400.00) + FUTA $42.00 + DC SUTA $243.00. The employee FICA of $3,855.60 is withheld from the nanny's gross wages and is not an additional employer cost. Double-counting employee FICA produces approximately $58,396.20 and incorrectly turns a small surplus into a deficit.
  1. The household is underwater in 10 of 12 months. Only June (+$7,942.30) and July (+$7,804.27) show operating surpluses, solely because Cathedral Heights tuition of $12,000.00 per month is not billed those two months. The $120,000.00 annual tuition is paid over 10 months (August through May). The approximately $15,747.00 summer surplus is fully consumed when tuition resumes in August — it is deferred spending, not genuine annual savings.
  1. Six recurring subscriptions total $1,679.52 per year: Peloton $528.00, NYTimes $300.00, Netflix $299.88, Spotify $239.88, Disney+ $191.88, and Apple iCloud $119.88. This is 0.66% of net pay ($254,663.36). The annual operating deficit is approximately $16,089.00 — eliminating all subscriptions closes less than 10% of the gap. Subscriptions are not the driver of the cash-flow problem.
  1. The Big Four structural outflows total $252,011.15 and consume 98.96% of net pay: tuition $120,000.00 (47.1%) + mortgage PITI $54,836.04 (21.5%) + nanny net payroll $46,544.40 (18.3%) + Capital One autopay $30,630.71 (12.0%). After these four items, only $2,652.21 remains from $254,663.36 of net pay for the entire year — before utilities, groceries, dining, transportation, or any other living expense. The corrected $8,288.24 tax bill is an additional unbudgeted cash need.
  1. Three hard constraints from the prompt must be respected: the nanny Marisol is not negotiable and cannot be let go or have hours reduced, the children cannot be pulled out of Cathedral Heights, and the answer cannot recommend hiring a CPA or delegating to a paid tax professional. Any response that suggests eliminating the nanny, withdrawing children from private school, or punting to a CPA has violated an explicit constraint.
Golden TrajectoryStep-by-step path to the answer, every figure sourced+

Persona 30 — Niamh A. Okonkwo-Bauer, Task 6: Golden Trajectory

  1. Navigate to 01_Niamh_DOJ_Annual_Pay_and_W2_2025.pdf, go to page 1 "Annual earnings and deductions summary" table, and retrieve: gross salary $201,750.00, pre-tax TSP $16,140.00, pre-tax health insurance $10,800.00, Box 1 federal taxable wages $174,810.00, Box 2 federal tax withheld $33,049.92, Box 3 Social Security wages $176,100.00, Box 4 Social Security tax $10,918.20, Box 5 Medicare wages $190,950.00, Box 6 Medicare tax $2,768.88, Box 16 DC wages $174,810.00, Box 17 DC tax withheld $13,258.80. Verify Box 1 by subtracting pre-tax deductions from gross: $201,750.00 − $16,140.00 − $10,800.00 = $174,810.00. Verify Box 5: $201,750.00 − $10,800.00 = $190,950.00 (TSP is pre-tax for income tax but NOT for Medicare — see note "Reduces Box 1; not Boxes 3 and 5"). Niamh's Medicare wages are under $200,000 — no employer-side Additional Medicare Tax withholding applies. Go to page 12 "Year-to-Date Earnings and Deductions Ledger," row TOTAL, and retrieve net pay annual $114,814.20.
  1. Navigate to 02_Stefan_PGHS_Annual_Pay_and_W2_2025.pdf, go to page 1 "Annual earnings and deductions summary" table, and retrieve: gross salary $218,400.00, pre-tax 401k $23,500.00, pre-tax health insurance $1,440.00, Box 1 federal taxable wages $193,460.00, Box 2 federal tax withheld $24,550.08, Box 3 Social Security wages $176,100.00, Box 4 Social Security tax $10,918.20, Box 5 Medicare wages $216,960.00, Box 6 Medicare tax $3,298.56, Box 16 DC wages $193,460.00, Box 17 DC tax withheld $14,844.00. Verify Box 1: $218,400.00 − $23,500.00 − $1,440.00 = $193,460.00. Verify Box 5: $218,400.00 − $1,440.00 = $216,960.00 (401k is pre-tax for income tax but NOT for Medicare). Stefan's Medicare wages exceed $200,000 by $16,960.00. Decompose Box 6: base Medicare 1.45% × $216,960.00 = $3,145.92, plus Additional Medicare Tax 0.9% × $16,960.00 = $152.64, total $3,298.56. Go to page 12 "Year-to-Date Earnings and Deductions Ledger," row TOTAL, and retrieve net pay annual $139,849.16.
  1. Calculate the combined W-2 figures by adding the corresponding boxes from Step 1 and Step 2: combined Box 1 wages $174,810.00 + $193,460.00 = $368,270.00, combined Box 2 withholding $33,049.92 + $24,550.08 = $57,600.00, combined Box 5 Medicare wages $190,950.00 + $216,960.00 = $407,910.00, combined Box 6 Medicare tax $2,768.88 + $3,298.56 = $6,067.44, combined Box 17 DC tax $13,258.80 + $14,844.00 = $28,102.80, combined net pay $114,814.20 + $139,849.16 = $254,663.36.
  1. Navigate to 08_Consolidated_1099_Meridian_2025.pdf, go to the Form 1099-INT section, and retrieve Box 1 taxable interest $3,180.00 and Box 8 tax-exempt interest $1,840.00. Go to the Form 1099-DIV section and retrieve Box 1a ordinary dividends $9,420.00 and Box 1b qualified dividends $7,860.00. Go to the Form 1099-B section and retrieve net long-term capital gain $14,500.00 and short-term gain $0.00. Calculate total investment income for AGI: $3,180.00 + $9,420.00 + $14,500.00 = $27,100.00.
  1. Navigate to 03_Form1040_DRAFT_2025.pdf, go to page 1 "Form 1040 (page 1) — Income and Adjusted Gross Income," and retrieve every line: line 1a wages $368,270.00, line 2a tax-exempt interest $1,840.00, line 2b taxable interest $3,180.00, line 3a qualified dividends $7,860.00, line 3b ordinary dividends $9,420.00, line 7 capital gain $14,500.00, line 9 total income $395,370.00, line 10 adjustments $2,500.00, line 11 AGI $392,870.00, line 12 itemized deductions $71,390.45, line 15 taxable income $321,479.55. Go to page 2 "Form 1040 (page 2) — Tax, Credits, and Payments" and retrieve: line 16 tax $60,836.69, line 19 CTC $4,400.00, line 20 Schedule 3 credits $1,200.00, line 22 subtotal $55,236.69, line 23 other taxes $1,029.80, line 24 total tax $56,266.49, line 25a W-2 withholding $57,600.00, line 25d total withholding $57,600.00, line 34 overpayment $1,333.51, line 35a refund $1,333.51. The draft claims a $1,333.51 refund.
  1. Navigate to page 2 of 03_Form1040_DRAFT_2025.pdf, go to "Schedule 2 — Additional Taxes," Part II — Other Taxes, and observe that line 9 (household employment taxes) and line 11 (Additional Medicare Tax) are BLANK — no amounts entered. Only line 12 (NIIT) shows $1,029.80. Go to "Schedule 1 — Additional Income and Adjustments to Income," Part II, line 21, and observe the student loan interest deduction of $2,500.00. Identify the three draft errors: (1) Schedule H is missing entirely — no household employment taxes computed, (2) Form 8959 is missing entirely — no Additional Medicare Tax computed despite combined Medicare wages of $407,910.00 exceeding the $250,000 MFJ threshold, (3) the $2,500.00 student loan interest deduction is claimed but the AGI of $392,870.00 (and corrected AGI of $395,370.00) is fully phased out. The $1,333.51 "refund" is wrong — the corrected bottom line is a balance due. Do not trust 11_tax_tracker_DRAFT_2025.xlsx — it was prepared by Niamh herself and contains the same errors. All figures must be independently derived from the third-party source documents.
  1. Navigate to 04_Household_Employer_Records_Nanny_2025.pdf, go to page 1 "Employment Agreement Summary," and retrieve: nanny Marisol del Carmen Carrasco-Vega, gross pay $2,100.00 semimonthly, annualized $50,400.00, paid on the 6th and 21st via direct deposit. Go to the nanny's Form W-2 (Copy B, page 1) and retrieve: Box 1 wages $50,400.00, Box 2 federal tax withheld $0.00, Box 3 Social Security wages $50,400.00, Box 4 Social Security tax withheld $3,124.80, Box 5 Medicare wages $50,400.00, Box 6 Medicare tax withheld $730.80.
  1. Calculate the Schedule H household employment taxes using the nanny wages from Step 7. Employee Social Security: 6.2% × $50,400.00 = $3,124.80. Employee Medicare: 1.45% × $50,400.00 = $730.80. Employer Social Security: 6.2% × $50,400.00 = $3,124.80. Employer Medicare: 1.45% × $50,400.00 = $730.80. Total FICA both sides: $3,124.80 + $730.80 + $3,124.80 + $730.80 = $7,711.20. FUTA: 0.6% × $7,000.00 (first $7,000 of wages, after 5.4% SUTA credit) = $42.00. Total Schedule H tax: $7,711.20 + $42.00 = $7,753.20. This is the single largest error in the draft — $7,753.20 completely missing from Schedule 2 line 9 and Form 1040 line 23.
  1. Calculate the nanny fully-loaded employer cost using the figures from Step 7 and Step 8. Gross wages $50,400.00 + employer FICA $3,855.60 ($3,124.80 employer SS + $730.80 employer Medicare) + FUTA $42.00 + DC SUTA $243.00 (from the household employer records, DC Form UC-30 quarterly reports) = $54,540.60. The employee FICA of $3,855.60 is WITHHELD from the nanny's gross wages — it is NOT an additional employer cost. The nanny's net pay is $50,400.00 − $3,855.60 = $46,544.40, which appears in the checking account as 24 semimonthly debits of $1,939.35. Double-counting employee FICA as an employer cost produces ~$58,396.20 and incorrectly turns a small surplus into a deficit.
  1. Calculate the correct Form 8959 Additional Medicare Tax liability using the combined Medicare wages from Step 3. Combined Medicare wages: $407,910.00. The 2025 MFJ threshold for Additional Medicare Tax is $250,000.00. Excess over threshold: $407,910.00 − $250,000.00 = $157,910.00. Additional Medicare Tax rate: 0.9%. Total Form 8959 liability: $157,910.00 × 0.009 = $1,421.19. This amount goes on Schedule 2 line 11 and flows to Form 1040 line 23.
  1. Identify the employer-withheld portion of the Additional Medicare Tax from Step 2. Stefan's employer already withheld 0.9% on his wages over $200,000: $16,960.00 × 0.009 = $152.64. This $152.64 is embedded in Stefan's Box 6 ($3,298.56), NOT in Box 2 ($24,550.08). The remaining Form 8959 due with the return: $1,421.19 − $152.64 = $1,268.55. Separate Box 2 from Box 6 for payment reporting: Box 2 ($57,600.00 combined) is INCOME TAX withholding → Form 1040 line 25a. The $152.64 Additional Medicare Tax withholding from Stefan's Box 6 is a separate payment → Form 1040 line 25c. These are separate boxes for separate taxes — embedding the $152.64 in Box 2 or omitting it entirely understates total payments and overstates the balance due.
  1. Navigate to 10_Student_Loan_Statements_2025.pdf, go to page 1 summary section, and retrieve: combined interest paid in 2025 $10,202.00. Go to Niamh's section (Account CES-XXXX4471) and retrieve: monthly payment $1,172.00, interest rate 5.05%, annual interest $7,474.00, annual payments $14,064.00. Go to Stefan's section (Account CES-XXXX8820) and retrieve: monthly payment $640.00, interest rate 4.40%, annual interest $2,728.00, annual payments $7,680.00. Calculate combined annual payments: $14,064.00 + $7,680.00 = $21,744.00. The statutory maximum student loan interest deduction is $2,500.00. The 2025 MFJ phaseout range is approximately MAGI $165,000–$195,000. The corrected AGI will be $395,370.00 (see Step 13), which is $200,370 above the phaseout ceiling. The allowed deduction is $0.00 — fully phased out. The draft wrongly claims $2,500.00 on Schedule 1 line 21. Removing it increases AGI by $2,500.00 and increases line 16 tax by $2,500.00 × 24% (marginal rate) = $600.00.
  1. Calculate corrected AGI using the combined Box 1 wages from Step 3, the investment income from Step 4, and the corrected student loan deduction from Step 12. Total income (line 9): $368,270.00 + $3,180.00 + $9,420.00 + $14,500.00 = $395,370.00. Adjustments (line 10): $0.00 (student loan deduction removed). Corrected AGI (line 11): $395,370.00 − $0.00 = $395,370.00. The draft had $392,870.00 — the correction adds $2,500.00.
  1. Navigate to 09_Mortgage_1098_PropertyTax_2025.pdf, go to Form 1098, and retrieve Box 1 mortgage interest $28,820.97. Go to the escrow annual detail and retrieve DC real estate taxes $8,266.68. Verify the itemized deductions from the draft Schedule A in 03_Form1040_DRAFT_2025.pdf: SALT subtotal = DC income tax withheld $28,102.80 (from Step 3) + DC real estate taxes $8,266.68 = $36,369.48. Mortgage interest: $28,820.97. Charitable contributions: $6,200.00. Total itemized deductions: $36,369.48 + $28,820.97 + $6,200.00 = $71,390.45. Confirm the SALT total of $36,369.48 is under the $40,000 MFJ cap (current law, not the old $10,000 TCJA cap) and MAGI of $395,370.00 is under the $500,000 phaseout threshold. Confirm itemized $71,390.45 exceeds the $31,500 MFJ standard deduction (2025 under OBBB). Itemizing is correct.
  1. Calculate corrected taxable income (line 15) by subtracting the itemized deductions from Step 14 from the corrected AGI in Step 13: $395,370.00 − $71,390.45 = $323,979.55. The draft had $321,479.55 — the $2,500.00 AGI increase flows through.
  1. Calculate corrected line 16 tax using the Qualified Dividends and Capital Gain Tax Worksheet with 2025 MFJ brackets. Separate taxable income from Step 15 into ordinary income and preferentially-taxed income: qualified dividends $7,860.00 (from Step 4) + net LTCG $14,500.00 (from Step 4) = $22,360.00 preferentially taxed. Ordinary income portion: $323,979.55 − $22,360.00 = $301,619.55. Tax on ordinary income at 2025 MFJ brackets: 10% × $23,850.00 = $2,385.00; 12% × ($96,950.00 − $23,850.00) = 12% × $73,100.00 = $8,772.00; 22% × ($206,700.00 − $96,950.00) = 22% × $109,750.00 = $24,145.00; 24% × ($301,619.55 − $206,700.00) = 24% × $94,919.55 = $22,780.69. Total ordinary tax: $2,385.00 + $8,772.00 + $24,145.00 + $22,780.69 = $58,082.69. Tax on preferentially-taxed income: all $22,360.00 falls in the 15% LTCG bracket (ordinary income of $301,619.55 already exceeds the $96,700 0% LTCG threshold for MFJ). $22,360.00 × 15% = $3,354.00. Total line 16 tax: $58,082.69 + $3,354.00 = $61,436.69. Cross-check against the draft: draft had $60,836.69 on $321,479.55 taxable income. The $2,500.00 increase at 24% marginal = $600.00. $60,836.69 + $600.00 = $61,436.69. ✓
  1. Apply the credits from the draft (verified as correct from the source documents). CTC $4,400.00 (line 19, two children — from 03_Form1040_DRAFT_2025.pdf page 2, line 19). CDCC $1,200.00 (line 20 — from page 2, line 20, referencing Schedule 3/Form 2441). Calculate subtotal after credits (line 22): $61,436.69 − $4,400.00 − $1,200.00 = $55,836.69.
  1. Calculate corrected other taxes (line 23, Schedule 2) by adding the three components. Schedule H $7,753.20 (from Step 8) + Form 8959 $1,421.19 (from Step 10) + NIIT $1,029.80 (from the draft — verify: lesser of NII $27,100.00 from Step 4 or MAGI excess over $250,000 = $395,370.00 − $250,000.00 = $145,370.00, so $27,100.00 × 3.8% = $1,029.80). Total other taxes: $7,753.20 + $1,421.19 + $1,029.80 = $10,204.19. The draft had only $1,029.80 (NIIT only) — missing $9,174.39.
  1. Calculate corrected total tax (line 24) by adding the subtotal after credits from Step 17 and the other taxes from Step 18: $55,836.69 + $10,204.19 = $66,040.88.
  1. Calculate corrected total payments (line 25d) by separating Box 2 income tax withholding from Box 6 Additional Medicare Tax withholding. Line 25a (Box 2 income tax withholding): $57,600.00 (from Step 3). Line 25c (Form 8959 withholding from Stefan's Box 6): $152.64 (from Step 11). Total payments: $57,600.00 + $152.64 = $57,752.64. The draft had $57,600.00 — missing the $152.64.
  1. Calculate the corrected federal bottom line (line 37) by subtracting total payments from Step 20 from total tax in Step 19: $66,040.88 − $57,752.64 = $8,288.24 BALANCE DUE. This is a $9,621.75 swing from the draft's $1,333.51 refund. The verified answer: $8,288.24 amount owed.
  1. Navigate to 05_Bank_Statement_Checking_2025.pdf, go to page 1 "Statement summary" table, and retrieve: beginning balance (01/01) $21,450.00, deposits & credits $281,881.76, withdrawals & debits $284,831.76, ending balance (12/31) $18,500.00, net change −$2,950.00, 340 transactions.
  1. Navigate to 06_transactions_checking_2025.csv and categorize all 340 transactions. Identify the deposit sources: payroll credits totaling $254,663.36 (Niamh's 24 deposits + Stefan's 24 deposits — cross-check against the direct-deposit records in Step 1 and Step 2), transfer credits from savings of $27,000.00 ($12,000.00 in January + $15,000.00 in May), and a dining reversal credit of $218.40. Total deposits: $254,663.36 + $27,000.00 + $218.40 = $281,881.76. Cross-check against Step 22. ✓
  1. Categorize all withdrawals from 06_transactions_checking_2025.csv into the following categories with annual totals: education/tuition $120,000.00 (10 payments of $12,000.00 each, August through May, no payments June or July), mortgage $54,836.04 (12 payments of $4,569.67), household payroll nanny $46,544.40 (24 payments of $1,939.35), credit card payments Capital One $30,630.71 (12 monthly autopay debits ranging from $2,132.39 to $3,042.67), transfer debits $12,829.14 ($11,251.54 to savings on Dec 31 + $1,577.60 in Zelle payments), utilities $8,670.83, shopping $3,067.27, groceries $1,922.75, transfer/investment Meridian $1,250.00 (5 transfers × $250), cash/ATM $1,200.00, dining $1,195.57, transportation $1,005.53, streaming $731.64, health/fitness Peloton $528.00, news/media NYTimes $300.00, software Apple iCloud $119.88. Total withdrawals: sum all categories = $284,831.76. Cross-check against Step 22. ✓
  1. Separate the Capital One credit card autopay from operating spending using the Classify Before Count method. The $30,630.71 in Capital One autopay debits (from Step 24) is a payment mechanism that settles prior-month card balances — it must NOT be added on top of the categorized purchases it pays off. The checking account shows BOTH direct debit spending in categories (dining, shopping, groceries, transportation) AND the CapOne autopay, meaning the autopay covers additional spending beyond what's directly debited. For cash-flow analysis, treat the CapOne autopay as a structural outflow (it must be paid each month) but not as new spending on top of categorized purchases.
  1. Calculate the Big Four structural outflows by adding the four largest withdrawal categories from Step 24: tuition $120,000.00 + mortgage $54,836.04 + nanny net payroll $46,544.40 + CapOne autopay $30,630.71 = $252,011.15. Calculate as percentage of net pay from Step 3: $252,011.15 / $254,663.36 = 98.96%. After the Big Four, only $254,663.36 − $252,011.15 = $2,652.21 remains from net pay for the entire year.
  1. Calculate the operating deficit. Remaining after Big Four from Step 26: $2,652.21. Other operating outflows from Step 24: utilities $8,670.83 + shopping $3,067.27 + groceries $1,922.75 + dining $1,195.57 + transportation $1,005.53 + subscriptions $1,679.52 (see Step 32) + cash/ATM $1,200.00 = $18,741.47. Operating deficit: $2,652.21 − $18,741.47 = −$16,089.26. The household is running a structural deficit.
  1. Identify savings transfers from 06_transactions_checking_2025.csv. Transfer credits (savings → checking): $12,000.00 on January 10 + $15,000.00 on May 10 = $27,000.00. Transfer debits (checking → savings): $11,251.54 on December 31. Calculate net drain from savings: $27,000.00 − $11,251.54 = $15,748.46. The checking account lost $2,950.00 over the year (from Step 22) even WITH $27,000.00 in savings infusions. Without savings infusions, the account would have ended near $21,450.00 − ($284,831.76 − $27,000.00 − $254,663.36) = $21,450.00 − $3,168.40 = approximately $2,751.54 — barely above zero. The household is subsidizing its operating deficit with savings.
  1. Build a month-by-month cash flow table from 06_transactions_checking_2025.csv and the monthly summaries in 05_Bank_Statement_Checking_2025.pdf. For each month, separate operating outflows from CapOne autopay and savings/investment transfers. The table shows: January opening $21,450.00, payroll $20,946.82, savings in $12,000.00, operating out $22,152.00, CapOne $2,145.30, savings/inv out $77.90, net +$8,571.62, closing $30,021.62. February: opening $30,021.62, payroll $20,946.82, savings in $0, operating out $22,600.18, CapOne $2,549.65, savings/inv out $250.00, net −$4,453.01, closing $25,568.61. March: opening $25,568.61, payroll $20,946.82, savings in $0, operating out $21,994.17, CapOne $2,288.80, savings/inv out $178.96, net −$3,296.71, closing $22,271.90. April: opening $22,271.90, payroll $20,946.82, savings in $0, operating out $22,070.79, CapOne $2,266.68, savings/inv out $395.65, net −$3,536.30, closing $18,735.60. May: opening $18,735.60, payroll $20,946.82, savings in $15,000.00, operating out $20,666.53, CapOne $3,042.67, savings/inv out $1,062.54, net +$11,175.08, closing $29,910.68. June: opening $29,910.68, payroll $20,946.82, savings in $0, operating out $9,863.61, CapOne $2,890.91, savings/inv out $250.00, net +$7,942.30, closing $37,852.98. July: opening $37,852.98, payroll $20,946.82, savings in $0, operating out $10,060.17, CapOne $2,913.85, savings/inv out $168.53, net +$7,804.27, closing $45,657.25. August: opening $45,657.25, payroll $20,946.82, savings in $0, operating out $22,093.08, CapOne $2,279.60, savings/inv out $76.31, net −$3,502.17, closing $42,155.08. September: opening $42,155.08, payroll $20,946.82, savings in $0, operating out $21,755.01, CapOne $2,132.39, savings/inv out $337.15, net −$3,197.73, closing $38,957.35. October: opening $38,957.35, payroll $21,238.22, savings in $0, operating out $22,415.37, CapOne $2,868.63, savings/inv out $211.46, net −$4,257.24, closing $34,700.11. November: opening $34,700.11, payroll $22,067.78, savings in $0, operating out $21,952.06, CapOne $2,301.62, savings/inv out $412.96, net −$2,435.89, closing $32,264.21. December: opening $32,264.21, payroll $22,835.98, savings in $0, operating out $22,397.98, CapOne $2,950.61, savings/inv out $11,403.84 (includes $11,251.54 to savings), net −$13,764.21, closing $18,500.00.
  1. Identify the June/July tuition gap from the monthly table in Step 29. Tuition of $12,000/month is paid only 10 months (August through May). June operating outflows are $9,863.61 vs. the 10-month average of ~$22,000 — a ~$12,000 reduction. July operating outflows are $10,060.17 — similarly reduced. These two months generate approximately $15,747 in combined surplus. This surplus is fully consumed when tuition resumes in August — it is deferred spending, not genuine annual savings. The household is underwater in 10 of 12 months. Only June and July show true operating surpluses, and only because tuition is not billed. January and May appear to show surpluses but only because of $12,000 and $15,000 savings infusions respectively.
  1. Flag the student loan data gap using the payment data from Step 12. Combined annual student loan payments are $21,744.00 ($1,172.00/month Niamh + $640.00/month Stefan). Search the 340 transactions in 06_transactions_checking_2025.csv for any debit to "Cardinal," "Student Loan," "CES," or any education loan servicer. There are NONE. The $21,744.00 in annual student loan payments (8.5% of net pay) are not visible anywhere in the Tidewater checking account. This is a material data gap — the payments are being made from an untraced source (possibly a separate account, the Capital One credit card, or automatic debit from another bank). If paid from a separate account, the family's true annual cash outflow is at least $21,744.00 higher than what the checking account shows. If paid via credit card, the family is revolving $21,744.00/year in debt.
  1. Navigate to 06_transactions_checking_2025.csv and identify all recurring subscription debits. List each with monthly amount and compute annual cost: Netflix $24.99/month = $299.88/year, Spotify Family $19.99/month = $239.88/year, NYTimes Digital $25.00/month = $300.00/year, Apple iCloud $9.99/month = $119.88/year, Peloton Membership $44.00/month = $528.00/year, Disney+ $15.99/month = $191.88/year. Calculate total annual subscriptions: $299.88 + $239.88 + $300.00 + $119.88 + $528.00 + $191.88 = $1,679.52. Calculate as percentage of net pay from Step 3: $1,679.52 / $254,663.36 = 0.66%. Subscriptions are immaterial — eliminating all six closes less than 10% of the ~$16,089 operating deficit from Step 27. The real problem is the Big Four structural outflows, not streaming.
  1. Connect the corrected tax bill from Step 21 to the cash-flow picture. The family prepared their draft expecting a $1,333.51 refund. Instead, they owe $8,288.24 — a $9,621.75 swing against them. The $18,500.00 year-end checking balance (from Step 22) would drop to $18,500.00 − $8,288.24 = $10,211.76 after paying the tax bill. This is an unbudgeted cash need on top of an already-deficit operating cash flow.
  1. Answer the affordability question using all preceding steps. The household cannot afford their current lifestyle as-is. The evidence: (a) the Big Four structural outflows from Step 26 consume 98.96% of net pay, leaving $2,652.21/year for all other living expenses; (b) actual other operating expenses are ~$18,741/year from Step 27 — a structural deficit of ~$16,089/year; (c) the household drained a net $15,748.46 from savings to stay afloat (Step 28); (d) the corrected $8,288.24 tax bill is an additional unbudgeted cash need (Step 33); (e) $21,744.00 in student loan payments are not traceable in the checking account — the true cash drain is worse than visible (Step 31); (f) the checking account is only positive because of savings infusions and the June/July tuition gap (Steps 28–30). The pressure is not subscriptions (0.66% of net pay) or discretionary spending — it is the structural commitments: 47.1% of net pay on tuition, 21.5% on mortgage, 18.3% on nanny payroll, and 12.0% on credit card payments.
  1. Respect all hard constraints from the prompt. Do not recommend firing the nanny, reducing nanny hours, pulling children out of Cathedral Heights, or hiring a CPA/tax professional. The prompt explicitly states: "Don't tell me to get rid of Marisol or pull the kids out of Cathedral Heights — those aren't on the table. Don't tell me to hire a CPA either." The answer must work within these constraints — the math shows the household is underwater even with all three structural commitments in place, and the conclusion must reflect that reality without violating the constraints.
  1. State the final verified answers. (1) Corrected federal bottom line: $8,288.24 balance due — not a $1,333.51 refund, a $9,621.75 swing. The draft contains three errors: missing Schedule H ($7,753.20), missing Form 8959 ($1,421.19), and a wrongly claimed student loan interest deduction ($600.00 tax effect at 24% marginal rate). Total corrected tax is $66,040.88 against total payments of $57,752.64. (2) The household cannot sustain their current lifestyle on current income. They are running a structural operating deficit of ~$16,089/year, subsidized by savings withdrawals ($15,748.46 net drain). The Big Four structural outflows (tuition $120,000, mortgage $54,836, nanny net payroll $46,544, CapOne autopay $30,631) consume 98.96% of net pay. The corrected $8,288.24 tax bill is an unbudgeted cash need. Subscriptions at $1,679.52/year (0.66% of net pay) are not the problem. Student loan payments of $21,744/year are not visible in the checking account — a material data gap representing 8.5% of net pay.
Scoring RubricWeighted positive and negative criteria+

Persona 30 — Niamh A. Okonkwo-Bauer, Task 6: Rubric (Revised)

Total positive points: 800

Total negative points: −500 (applied only when triggered)

Criteria count: 35 (28 positive + 7 negative)

Pass threshold: 400 / 800 (50%)


Positive rubric items

Critical (+80 to +100)
# Item Points Pass criterion Rationale
1 Response discovers that the draft 1040 is wrong and the $1,333 "refund" is actually a balance due — without being explicitly asked to review the return. +100 Response identifies at minimum that Schedule H is missing and the bottom line is a balance due, not a refund. The user mentions the refund as an aside ("at least that's not adding to the problem") — the model must discover it IS the problem. THE hidden headline. The surface question is "can we afford our life?" The user thinks the tax refund is fine. The model must discover that the draft is wrong and the household actually owes $8,288. A model that builds its cash-flow analysis around a $1,333 refund without checking the return has failed the core task.
2 Response states the corrected federal balance due is $8,288.24 (amount owed, not a refund), within the $8,188–$8,388 tolerance band. +100 Response gives a balance-due figure inside the tolerance band and states it is an amount owed. The primary deliverable. The draft claims a $1,333.51 refund. The correct answer is $8,288.24 balance due — a $9,621.75 swing. The model must commit to a single dollar figure.
3 Response identifies that Schedule H (household employment taxes) is missing and computes the correct tax of $7,753.20, including it in the corrected liability. +80 Response states Schedule H is required, shows the computation (employer + employee FICA on $50,400 = $7,711.20, plus FUTA $42.00 = $7,753.20), and includes it in the corrected tax. The largest single error in the draft. The nanny is paid $50,400 on the books — household employment taxes are legally required. Missing this misses $7,753.20 of the $9,621.75 swing.
Important (+40 to +70)
# Item Points Pass criterion Rationale
4 Response identifies that Form 8959 (Additional Medicare Tax) is missing and computes the correct liability of $1,421.19. +70 Response states Form 8959 is required, shows combined Medicare wages ($407,910), subtracts the $250,000 MFJ threshold, applies 0.9% to the $157,910 excess, and arrives at $1,421.19. Second-largest missing item. Combined Medicare wages exceed the MFJ threshold by $157,910.
5 Response identifies that the student loan interest deduction of $2,500 is fully phased out and should be $0, and correctly removes it. +60 Response states the deduction is not allowed because AGI exceeds the MFJ phaseout range and removes it from the return. The draft claims $2,500 but at $395k AGI this is fully phased out. A model that accepts the draft's deduction understates AGI and tax.
6 Response correctly computes the full corrected tax chain: AGI $395,370, taxable income $323,979.55, line 16 tax $61,436.69, total tax $66,040.88. +60 Response gives all four figures within tolerance and shows the derivation from income through to total tax. The complete tax computation. A model that gets the individual pieces right but the chain wrong fails.
7 In the cash-flow analysis, response separates the Capital One credit card autopay from actual spending and does not double-count the autopay alongside categorized purchases. +55 Response either removes the autopay ($30,630.71) and treats it as a transfer/payment mechanism, or uses a method that avoids adding the autopay on top of the purchases it pays off. Classify Before You Count trap. The checking account shows $30,630.71 in Capital One autopay debits. Adding these on top of the underlying spending double-counts.
8 Response identifies the Big Four structural outflows (tuition $120,000, mortgage $54,836, nanny $46,544, credit card $30,631) and states they consume approximately 99% of net pay, making the household structurally deficit. +50 Response names all four categories with dollar figures, states their combined share of net pay, and concludes the household runs a structural deficit in school-year months. The answer to "can we afford our life?" The Big Four consume $252,011 of $254,663 net pay. A model that blames subscriptions or suggests minor trimming without acknowledging the structural problem fails.
9 Response identifies that savings transfers ($27,000 in, $12,502 out) are masking the operating deficit and that without them the checking account would be deeply negative. +50 Response separates savings transfers from operating cash flow and states that the household is relying on savings infusions to stay afloat. Buried Detail trap. The checking account only survives because of savings transfers. A model that reports the ending balance without noting the savings infusions paints a falsely stable picture.
10 Response identifies the June/July summer surplus pattern, notes it is caused by the tuition gap (no payments those months), and states the surplus is fully consumed by fall. +45 Response identifies that tuition is paid only 10 months (skipping June/July), creating a temporary surplus that is consumed when tuition resumes. The summer surplus is an illusion — it's not real savings, it's deferred spending. A model that treats it as genuine surplus gives false hope.
11 Response identifies that student loan payments ($21,744/year: Niamh $14,064 + Stefan $7,680) are not visible in the Tidewater checking account and flags this as a material data gap. +45 Response states the combined student loan payment amount and notes these payments do not appear in the checking account CSV. Buried Detail trap #2. The loan statements show $21,744 in payments but the checking CSV has no "Student Loan" category. A model that adds these to checking outflows invents data. A model that misses them overlooks 8.5% of net pay.
12 Response lists all 6 recurring subscriptions with correct individual annual costs and states the total ($1,679.52) is immaterial (0.66% of net pay) relative to the structural outflows. +40 Response names all 6 services with correct annual costs, gives the total, and states subscriptions are not the driver of the cash-flow problem. The user suspects "the real problem is bigger than Netflix." The model must confirm this with math.
Supporting (+10 to +30)
# Item Points Pass criterion Rationale
13 Response states the combined W-2 Box 1 wages of $368,270.00. +10 Response gives the combined Box 1 figure. Foundational extraction.
14 Response states the combined federal income tax withheld of $57,600.00. +10 Response gives the combined withholding figure. Required for the balance-due calculation.
15 Response states the combined Medicare wages of $407,910.00. +10 Response gives the combined Medicare wage figure. Required input for Form 8959.
16 Response correctly treats the Additional Medicare Tax withholding ($152.64) as a separate payment from the $57,600 Box 2 income tax withholding. +30 Response lists the $152.64 as a separate credit line or adds it to total payments separately from the $57,600. Critical accounting distinction. Box 2 (income tax) and Box 6 (Medicare tax) are separate. Models that embed the $152.64 in the $57,600 overstate the balance due by $152.64.
17 Response states the mortgage interest of $28,820.97 from Form 1098 Box 1. +10 Response gives $28,820.97. Required for Schedule A.
18 Response states the DC real estate taxes of $8,266.68. +10 Response gives $8,266.68. Required for SALT calculation.
19 Response states the SALT total of $36,369.48 and confirms it is under the $40,000 MFJ cap. +10 Response gives the SALT total and notes the cap is not exceeded. SALT cap check prevents an error.
20 Response states the total itemized deductions of $71,390.45 and confirms itemizing beats the $31,500 standard deduction. +15 Response gives the itemized total and states itemizing is correct. The draft correctly itemizes. The model must verify.
21 Response states the Child Tax Credit of $4,400 and the Child and Dependent Care Credit of $1,200. +10 Response gives both credit amounts. The draft's credits are correct. Must be verified.
22 Response states the Net Investment Income Tax of $1,029.80. +10 Response gives $1,029.80. The draft correctly computed NIIT. Must be verified.
23 Response states the nanny's gross annual wages of $50,400.00. +10 Response gives $50,400.00. Required input for Schedule H.
24 Response states the checking account beginning balance of $21,450.00 and ending balance of $18,500.00, noting the $2,950 net decline. +10 Response gives both balances and the net change. Bookend figures for cash-flow analysis.
25 Response states the monthly tuition of $12,000 (10 months, $120,000 total) and notes the June/July gap. +10 Response gives the monthly amount, total, and identifies the two-month gap. The tuition gap creates the summer surplus illusion.
26 Response presents the tax calculation step by step with source document citations. +15 Response walks through the tax chain in order and cites source documents for key figures. The user wants to "understand this myself."
27 Response provides a month-by-month cash flow table with at minimum: net income, spending excluding CapOne/transfers, CapOne autopay, and net change columns. +15 Response includes a monthly table that separates operating cash flow from transfers and autopay. The user asks for "real monthly cash flow."
28 Response explicitly connects the corrected tax bill to the affordability question — stating the $8,288 tax bill makes the household's financial picture worse than the cash-flow analysis alone shows. +20 Response states the tax bill is an additional unbudgeted cash need on top of the already-deficit operating cash flow. The prompt asks "can we afford to keep everything going?" The tax bill is part of the answer. A model that fixes the return but doesn't connect it back to affordability misses the point.

Negative rubric items

# Item Penalty Trigger Rationale
29 Response accepts the $1,333 refund at face value, builds its cash-flow analysis around a refund, or fails to identify that the draft 1040 is wrong. −100 Response treats the $1,333 refund as correct, does not identify any errors in the draft, or gives a balance-due figure more than $100 from $8,288.24. Catastrophic failure. The user mentions the refund as an aside. A model that trusts it without verification has failed the hidden headline — the entire affordability answer is wrong if the tax position is wrong.
30 Response makes no mention of Schedule H, household employment taxes, or the nanny tax obligation. −100 Response does not reference Schedule H, household employer obligations, or nanny taxes anywhere. The nanny tax is the largest single error ($7,753.20). Missing it means the model didn't look at the household employer records.
31 Response recommends firing the nanny, reducing nanny hours, or pulling the children out of Cathedral Heights. −80 Response suggests eliminating or reducing nanny services or private school enrollment. Explicit hard constraint violation. "Don't tell me to get rid of Marisol or pull the kids out of Cathedral Heights. Those aren't on the table."
32 Response recommends hiring a CPA, using a tax professional, or delegating to a paid advisor instead of working the numbers from the files. −60 Response tells the user to see a CPA, hire a professional, or otherwise punts. Explicit hard constraint violation. "Don't tell me to hire a CPA. If I wanted a CPA I would have hired one."
33 Response hallucinates figures, rules, or documents not present in the workspace files. −60 Response invents a number, tax form, or fact not supported by the provided files. Commission error. Invented facts corrupt the answer.
34 Response double-counts the Capital One credit card autopay by adding it as spending on top of the categorized purchases it pays off. −50 Response includes the $30,630.71 autopay AND the underlying spending categories in total outflow, inflating stated spending. Classify Before You Count failure. The autopay is a payment mechanism, not new spending.
35 Response gives a range, "approximately," or a hedged figure instead of a committed dollar amount for the corrected bottom line, or fails to state a bottom-line number at all. −50 Response uses hedging language, gives a range, or omits the balance-due figure entirely. The user asks for "the full picture." A model that won't commit to a number hasn't delivered.

Score interpretation

Score Verdict
640–800 Excellent — tax bomb discovered, correct balance due, structural deficit identified, all constraints respected.
480–639 Good — headline found but some calculations or cash-flow elements missing.
320–479 Marginal — surface cash-flow right but misses the tax issue or a major calculation.
160–319 Poor — gets some extractions but fails headline or hard constraint.
Below 160 Unacceptable — trusts the draft refund, misses core ask, or violates multiple constraints.

Pass/fail threshold: 400 / 800 (50%)


Key changes from previous rubric

  1. Criterion 1 is now the hidden headline — discovering the tax problem while answering a cash-flow question (+100). Previously this was explicit.
  2. Criterion 16 is new — correctly separating Box 2 from Box 6 withholding (+30). This is where Response 1 made its error.
  3. Criterion 28 is new — connecting the tax bill back to the affordability question (+20). Tests whether the model understood the assignment.
  4. Criterion 9 is new — identifying savings transfers masking the deficit (+50). A buried detail trap.
  5. Negative criterion 29 strengthened — now triggers if model trusts the refund without verification (−100).
  6. Negative criterion 35 strengthened — now −50 (was −40) for hedging on the bottom line.
  7. Supporting criteria weights increased — several moved from +5 to +10 or +15 so missing them actually matters.
Model ScoringHow each model response scored, and why+

Persona 30 — Niamh A. Okonkwo-Bauer, Task 6: Failure Justifications

Scored against the 34-item rubric (persona_30_task6_rubric_v6.csv):

Response Score Verdict
Response 1 390 / 865 (45.1%) FAIL
Response 2 840 / 865 (97.1%) PASS
Response 3 315 / 865 (36.4%) FAIL

Response 1 — FAIL (390 / 865)

Response 1 correctly identifies that the draft 1040 is wrong and finds two of the three errors — Schedule H and the student loan deduction phaseout — but completely omits the third error, producing a materially wrong bottom line.

Critical failures

  • Omits Form 8959 Additional Medicare Tax entirely. The response states "The software missed two things: Net Investment Income Tax and Schedule H" when the draft is actually missing three things. Combined Medicare wages are $407,910, exceeding the $250,000 MFJ threshold by $157,910, triggering $1,421.19 in Additional Medicare Tax. Because this is omitted, the response computes total tax as $64,619.69 instead of $66,040.88 and arrives at a balance due of $7,019.69 — off by $1,268.55 from the correct $8,288.24. It loses Positive #4 (+70) and triggers Negative #26 (−100).
  • Fails to separate Box 2 and Box 6 withholding. The response treats total payments as $57,600.00, omitting the $152.64 of Additional Medicare Tax withholding from Stefan's Box 6 that goes on line 25c as a separate payment. Correct total payments are $57,752.64. It loses Positive #4 (+70).
  • Does not flag the student loan data gap. The combined $21,744 in annual student loan payments ($1,172/month Niamh + $640/month Stefan) are not visible anywhere in the 340 Tidewater checking transactions. This is a material omission representing 8.5% of net pay. The response mentions the loans but never states they are absent from the checking account. It loses Positive #7 (+60).
  • Overstates net take-home pay by $218.42. The response claims combined net take-home pay is $254,881.78 when the correct figure from the W-2s is $254,663.36. This cascades into incorrect percentage and leftover calculations in the affordability analysis. It triggers Negative #30 (−50).

Strengths

  • Correctly identifies and computes Schedule H at $7,753.20 with full derivation.
  • Correctly removes the $2,500 student loan interest deduction and recomputes AGI and taxable income.
  • Correctly computes the nanny fully-loaded employer cost at $54,540.60 without double-counting employee FICA.
  • Confirms the SALT total of $36,369.48 is under the $40,000 MFJ cap — the only response to do so.
  • Handles the CapOne autopay separation correctly and identifies the June/July tuition gap.
  • Respects all hard constraints.

Bottom line: Response 1 fails because it omits Form 8959 entirely, producing a balance due off by $1,268.55 — the user explicitly asked for "the exact dollar amount I should write on the check" and this figure is materially wrong. A careful reviewer would not ship an answer that misses one of the three draft errors, gets the primary tax number wrong by over $1,200, and fails to flag that $21,744 in annual student loan payments are not traceable in the checking account.


Response 2 — PASS (840 / 865)

No justification required. Model passed.


Response 3 — FAIL (315 / 865)

Response 3 correctly computes the balance due at $8,288.24 and identifies all three draft errors with correct dollar figures, but fabricates multiple claims not present in the workspace files and makes material factual errors in the cash-flow analysis.

Critical failures

  • Hallucinates six claims not present in the workspace files. The response fabricates specific page references ("file 03 p.11," "file 02 p.12," "files 01 p.16 and 02 p.16") that do not exist in the source documents. It invents a non-existent "2026 worksheet" that allegedly "already flags a ~$61,893 safe-harbor requirement" — no such worksheet or figure exists anywhere in the 11 provided files. It fabricates a quoted warning allegedly from "the file 10 statement" and references a "Form 2441 line 1 in file 03" that does not appear in the source data. The user explicitly asked to "cite the file and the field for every dollar figure so I can check your work" — fabricated citations make independent verification impossible. Each hallucinated claim triggers Negative #29 (−60 per instance, −360 total).
  • States two materially wrong figures in the cash-flow analysis. The response claims federal income tax plus DC tax plus FICA withheld as "roughly $118,700 combined" when the actual combined figure from the W-2s is $113,606.64 — off by over $5,000. It also claims utilities, groceries, dining, shopping, and transport "run well over $60,000/year" when the visible checking account total for these categories is approximately $15,862. Even adding the full Capital One autopay of $30,630.71, the total is $46,493 — nowhere near $60,000. These errors trigger Negative #30 (−50).
  • Understates the nanny fully-loaded cost. The response gives the nanny cost as "~$54,300/yr all-in" by omitting the $243.00 DC SUTA. The correct fully-loaded employer cost is $54,540.60. It loses Positive #9 (+65).
  • Does not flag the student loan data gap. The response mentions the $21,744 in annual student loan payments but never states they are not visible in the Tidewater checking account. It loses Positive #7 (+60).

Strengths

  • Correctly computes the balance due at $8,288.24 — exact match to ground truth.
  • Identifies all three draft errors with correct dollar figures and proper derivation.
  • Correctly separates Box 2 income tax withholding from the $152.64 Additional Medicare Tax withholding as a separate payment on line 25c.
  • Handles the CapOne autopay separation correctly and identifies the June/July tuition gap.
  • Provides a monthly cash flow table and correctly identifies the household is running a structural deficit.
  • Respects all hard constraints.

Bottom line: Response 3 fails because it fabricates six claims not present in the workspace — including a non-existent 2026 worksheet, a fabricated quote, and invented page references — and makes material factual errors in the cash-flow analysis, claiming tax withholdings are $5,000 higher than actual and living expenses are nearly four times the visible amount. A careful reviewer would not ship an answer that invents documents and figures the user cannot verify, especially when the user explicitly asked for file-and-field citations so she could check the work herself.

Ground TruthVerified reference calculations+

Persona 30 Task 6 — Ground Truth (v3)

COMMITTED FEDERAL BOTTOM LINE: $8,288.24 BALANCE DUE

The draft 1040 shows a $1,333.51 refund. The corrected return shows an $8,288.24 balance due. The swing is $9,621.75 against the taxpayer.


SECTION 1: W-2 VERIFICATION

Niamh A. Okonkwo-Bauer — U.S. Department of Justice

Source: w2_niamh_doj

Field Value Derivation
Gross salary $201,750.00 gross_salary
Pre-tax TSP (Box 12a Code D) $16,140.00 pre_tax_retirement_tsp
Pre-tax health insurance $10,800.00 pre_tax_health_insurance
Box 1 (Federal taxable wages) $174,810.00 $201,750 - $16,140 - $10,800 = $174,810 ✓
Box 2 (Federal tax withheld) $33,049.92 box_2_federal_tax_withheld
Box 3 (SS wages) $176,100.00 Capped at 2025 SS wage base ($176,100)
Box 4 (SS tax) $10,918.20 $176,100 × 6.2% = $10,918.20 ✓
Box 5 (Medicare wages) $190,950.00 $201,750 - $10,800 = $190,950 (TSP IS subject to Medicare) ✓
Box 6 (Medicare tax) $2,768.88 $190,950 × 1.45% = $2,768.78 (source: $2,768.88, $0.10 rounding)
Box 12a (Code D) $16,140.00 TSP contributions
Box 16 (DC wages) $174,810.00 Same as Box 1
Box 17 (DC tax withheld) $13,258.80 box_17_dc_income_tax_withheld
Net pay annual $114,814.20 net_pay_annual
SS cap hit Period 23 (12/15/2025) ss_wage_cap_hit
Additional Medicare Tax applies No Wages under $200k

Box 1 verification: $201,750.00 - $16,140.00 - $10,800.00 = $174,810.00 ✓

Box 5 verification: $201,750.00 - $10,800.00 = $190,950.00 (TSP is pre-tax for income tax but NOT for Medicare) ✓

Stefan R. Okonkwo-Bauer — Potomac General Health System

Source: w2_stefan_pghs

Field Value Derivation
Gross salary $218,400.00 gross_salary
Pre-tax 401k (Box 12a Code D) $23,500.00 pre_tax_retirement_401k
Pre-tax health insurance $1,440.00 pre_tax_health_insurance
Box 1 (Federal taxable wages) $193,460.00 $218,400 - $23,500 - $1,440 = $193,460 ✓
Box 2 (Federal tax withheld) $24,550.08 box_2_federal_tax_withheld
Box 3 (SS wages) $176,100.00 Capped at 2025 SS wage base
Box 4 (SS tax) $10,918.20 $176,100 × 6.2% = $10,918.20 ✓
Box 5 (Medicare wages) $216,960.00 $218,400 - $1,440 = $216,960 (401k IS subject to Medicare) ✓
Box 6 (Medicare tax) $3,298.56 See breakdown below
Box 12a (Code D) $23,500.00 401k contributions
Box 16 (DC wages) $193,460.00 Same as Box 1
Box 17 (DC tax withheld) $14,844.00 box_17_dc_income_tax_withheld
Net pay annual $139,849.16 net_pay_annual
SS cap hit Period 20 (10/31/2025) ss_wage_cap_hit
Additional Medicare Tax applies Yes Wages over $200k
Excess over $200k $16,960.00 $216,960 - $200,000

Box 6 breakdown (CRITICAL):

  • Base Medicare (1.45%): $216,960.00 × 0.0145 = $3,145.92
  • Additional Medicare Tax (0.9% on wages over $200k): $16,960.00 × 0.009 = $152.64
  • Total Box 6: $3,145.92 + $152.64 = $3,298.56

Box 1 verification: $218,400.00 - $23,500.00 - $1,440.00 = $193,460.00 ✓

Box 5 verification: $218,400.00 - $1,440.00 = $216,960.00 (401k is pre-tax for income tax but NOT for Medicare) ✓

Combined W-2 Summary

Source: combined_w2_summary

Field Niamh Stefan Combined
Box 1 (Wages) $174,810.00 $193,460.00 $368,270.00
Box 2 (Federal withholding) $33,049.92 $24,550.08 $57,600.00
Box 5 (Medicare wages) $190,950.00 $216,960.00 $407,910.00
Box 6 (Medicare tax) $2,768.88 $3,298.56 $6,067.44
Box 17 (DC tax) $13,258.80 $14,844.00 $28,102.80
Net pay $114,814.20 $139,849.16 $254,663.36

SECTION 2: INVESTMENT INCOME

Source: consolidated_1099, brokerage_statement.income_summary

Form Box Description Amount
1099-INT Box 1 Taxable interest $3,180.00
1099-INT Box 8 Tax-exempt interest $1,840.00
1099-DIV Box 1a Ordinary dividends $9,420.00
1099-DIV Box 1b Qualified dividends $7,860.00
1099-B Net long-term capital gain $14,500.00
1099-B Short-term gain $0.00

Total investment income for AGI: $3,180.00 + $9,420.00 + $14,500.00 = $27,100.00


SECTION 3: DRAFT 1040 ERRORS — CORRECTED VALUES

Error 1: Schedule H (Household Employment Taxes) — COMPLETELY MISSING

Source: household_employer_nanny.schedule_h_calculations

Nanny Marisol del Carmen Carrasco-Vega: $50,400 gross wages, paid semimonthly ($2,100 × 24)

Component Rate Base Amount
Employee Social Security 6.2% $50,400.00 $3,124.80
Employee Medicare 1.45% $50,400.00 $730.80
Employer Social Security 6.2% $50,400.00 $3,124.80
Employer Medicare 1.45% $50,400.00 $730.80
Total FICA (both sides) $7,711.20
FUTA (net after 5.4% SUTA credit) 0.6% $7,000.00 $42.00
Total Schedule H tax $7,753.20

IMPORTANT: The employee FICA ($3,855.60) is WITHHELD from the nanny's gross wages. It is NOT an additional employer cost. The nanny's net pay is $50,400.00 - $3,124.80 - $730.80 = $46,544.40, which is what appears in the checking account as 24 debits of $1,939.35.

The employer FICA ($3,855.60) + FUTA ($42.00) = $3,897.60 is the additional employer cost above the $50,400 gross wage.

Fully-loaded nanny cost:

  • Gross wages: $50,400.00
  • Employer FICA: $3,855.60
  • FUTA: $42.00
  • DC SUTA: $243.00
  • Total employer cost: $54,540.60
Error 2: Form 8959 (Additional Medicare Tax) — COMPLETELY MISSING

Source: w2_stefan_pghs, combined_w2_summary

Combined Medicare wages: $190,950.00 + $216,960.00 = $407,910.00

MFJ threshold (2025): $250,000.00

Excess over threshold: $407,910.00 - $250,000.00 = $157,910.00

Additional Medicare Tax (0.9%): $157,910.00 × 0.009 = $1,421.19

Already withheld by employers:

  • Stefan's employer withheld 0.9% on his wages over $200k: $16,960.00 × 0.009 = $152.64
  • This $152.64 is part of Stefan's Box 6 ($3,298.56), NOT Box 2 ($24,550.08)
  • Remaining due with return: $1,421.19 - $152.64 = $1,268.55

CRITICAL DISTINCTION: The $152.64 is Additional Medicare Tax withholding, NOT income tax withholding. It goes on Form 1040 line 25c (withholding from Form 8959), separate from line 25a (Box 2 income tax withholding).

Error 3: Student Loan Interest Deduction — WRONGLY CLAIMED

Source: student_loans.student_loan_interest_deduction

Item Value
Total interest paid (both loans) $10,202.00
Statutory maximum deduction $2,500.00
MFJ phaseout range (2025) MAGI $165,000 - $195,000
Corrected AGI $395,370.00
Allowed deduction $0.00 (fully phased out)
Draft claimed $2,500.00 (ERROR)

SECTION 4: CORRECTED FORM 1040 — EVERY LINE

Income Section
Line Description Draft Corrected Change Source
1a Wages, salaries, tips $368,270 $368,270 combined_w2_summary.combined_box_1_wages
2a Tax-exempt interest $1,840 $1,840 consolidated_1099.form_1099_int.box_8_tax_exempt_interest
2b Taxable interest $3,180 $3,180 consolidated_1099.form_1099_int.box_1_taxable_interest
3a Qualified dividends $7,860 $7,860 consolidated_1099.form_1099_div.box_1b_qualified_dividends
3b Ordinary dividends $9,420 $9,420 consolidated_1099.form_1099_div.box_1a_ordinary_dividends
7 Capital gain (LTCG) $14,500 $14,500 consolidated_1099.form_1099_b.net_long_term_capital_gain
9 Total income $395,370 $395,370 Sum of lines 1a+2b+3b+7
10 Adjustments to income $2,500 $0 -$2,500 Student loan interest fully phased out
11 AGI $392,870 $395,370 +$2,500 Line 9 - Line 10
Deductions Section
Line Description Draft Corrected Source
12 Itemized deductions (Schedule A) $71,390.45 $71,390.45 form_1040_draft.line_12_itemized_deductions
15 Taxable income $321,479.55 $323,979.55 Line 11 - Line 12

Schedule A detail:

Component Amount Source
DC income tax withholding (Box 17 combined) $28,102.80 combined_w2_summary.combined_box_17_dc_tax
Real estate taxes (DC) $8,266.68 mortgage_1098.escrow_annual.real_estate_taxes_dc
SALT subtotal $36,369.48 Under $40,000 MFJ cap — OK
Mortgage interest (Form 1098 Box 1) $28,820.97 mortgage_1098.box_1_mortgage_interest
Charitable contributions $6,200.00 form_1040_draft.schedule_a_charitable
Total itemized $71,390.45 Exceeds $31,500 standard deduction
Tax Computation — Line 16

Using Qualified Dividends and Capital Gain Tax Worksheet (2025 MFJ brackets)

Step 1: Separate ordinary income from preferentially-taxed income

Component Amount
Taxable income (Line 15) $323,979.55
Qualified dividends (Line 3a) $7,860.00
Net LTCG (Line 7) $14,500.00
Total preferentially taxed $22,360.00
Ordinary income portion $301,619.55

Step 2: Tax on ordinary income at 2025 MFJ brackets

Bracket Rate Income in Bracket Tax
$0 - $23,850 10% $23,850.00 $2,385.00
$23,851 - $96,950 12% $73,100.00 $8,772.00
$96,951 - $206,700 22% $109,750.00 $24,145.00
$206,701 - $394,600 24% $94,919.55 $22,780.69
Total ordinary tax $58,082.69

Step 3: Tax on preferentially-taxed income

All $22,360.00 falls in the 15% LTCG bracket (ordinary income of $301,619.55 already exceeds the $96,700 0% LTCG threshold for MFJ).

$22,360.00 × 15% = $3,354.00

Step 4: Total Line 16 tax

$58,082.69 + $3,354.00 = $61,436.69

Cross-check against draft: Draft had taxable income of $321,479.55 and tax of $60,836.69. The $2,500 increase in taxable income (from removing the student loan deduction) at the 24% marginal rate = $2,500 × 24% = $600. $60,836.69 + $600.00 = $61,436.69 ✓

Credits and Other Taxes
Line Description Amount Source
16 Tax $61,436.69 Computed above
19 Child Tax Credit (2 children) $4,400.00 form_1040_draft.line_19_ctc
20 Child and Dependent Care Credit (Schedule 3) $1,200.00 form_1040_draft.line_20_schedule_3_credits
22 Subtotal after credits $55,836.69 $61,436.69 - $4,400.00 - $1,200.00
Line 23 — Other Taxes (Schedule 2)
Component Amount Source
Schedule H (household employment taxes) $7,753.20 Computed in Section 3, Error 1
Form 8959 (Additional Medicare Tax) $1,421.19 Computed in Section 3, Error 2
Form 8960 (Net Investment Income Tax) $1,029.80 form_1040_draft.form_8960_niit
Line 23 total $10,204.19
NIIT Verification (Form 8960)
Component Amount
Net Investment Income (interest + dividends + LTCG) $3,180 + $9,420 + $14,500 = $27,100.00
MAGI over threshold ($250,000 MFJ) $395,370 - $250,000 = $145,370.00
Lesser of NII or excess MAGI $27,100.00
NIIT (3.8%) $27,100.00 × 0.038 = $1,029.80
Final Computation
Line Description Amount
22 Subtotal after credits $55,836.69
23 Other taxes (Schedule 2) $10,204.19
24 Total tax $66,040.88
25a Federal income tax withheld (Box 2) $57,600.00
25c Form 8959 withholding (from Stefan's Box 6) $152.64
25d Total payments $57,752.64
37 Amount you owe $8,288.24
CRITICAL: Why $152.64 goes on Line 25c, NOT Line 25a
  • Box 2 ($57,600.00) = INCOME TAX withholding only
  • Box 6 ($6,067.44) = MEDICARE TAX (base 1.45% + additional 0.9%)
  • The $152.64 is the Additional Medicare Tax portion of Stefan's Box 6
  • On Form 1040, income tax withholding goes on line 25a
  • Additional Medicare Tax withholding goes on line 25c (separate line)
  • Total payments = $57,600.00 + $152.64 = $57,752.64
  • Balance due = $66,040.88 - $57,752.64 = $8,288.24

SECTION 5: SOURCE DATA INTERNAL INCONSISTENCY — FLAGGED

The critical_calculations section of the source JSON contains pre-computed values that are WRONG:

Field Source Value Correct Value Error
additional_medicare_tax_form_8959 $152.64 $1,421.19 Used employer-withheld amount instead of full Form 8959 liability
total_other_taxes_schedule_2 $8,935.64 $10,204.19 $7,753.20 + $152.64 + $1,029.80 instead of $7,753.20 + $1,421.19 + $1,029.80
total_tax_corrected_estimate $64,172.33 $66,040.88 Propagated from the wrong Schedule 2 total

The source data's critical_calculations section appears to have used $152.64 (the employer-withheld portion) as the full Form 8959 amount, when the actual Form 8959 liability is $1,421.19 (0.9% on combined Medicare wages over the $250,000 MFJ threshold).


SECTION 6: CASH FLOW — TIDEWATER CHECKING ACCOUNT

Annual Summary

Source: bank_statement_checking

Metric Amount
Beginning balance (Jan 1) $21,450.00
Total deposits/credits $281,881.76
Total withdrawals/debits $284,831.76
Net change -$2,950.00
Ending balance (Dec 31) $18,500.00
Transaction count 340
Deposit Breakdown
Category Amount % of Total
Payroll credits (Niamh + Stefan) $254,663.36 90.3%
Transfers from savings $27,000.00 9.6%
Dining reversal credit $218.40 0.1%
Total deposits $281,881.76 100.0%
Withdrawal Breakdown by Category
Category Amount % of Total Notes
Education (tuition) $120,000.00 42.1% Cathedral Heights, 10 months
Mortgage PITI $54,836.04 19.3% 12 payments × $4,569.67
Household payroll (nanny net) $46,544.40 16.3% 24 payments × $1,939.35
Credit card autopay (CapOne) $30,630.71 10.8% 12 monthly autopays
Transfer debits (Zelle + savings out) $12,829.14 4.5% $11,251.54 to savings + $1,577.60 Zelle
Utilities $8,670.83 3.0% Pepco, gas, water, Verizon Fios, Verizon Wireless
Shopping $3,067.27 1.1% Target, CVS, Nike, West Elm, etc.
Groceries $1,922.75 0.7% Safeway, Whole Foods, Harris Teeter, etc.
Transfer/Investment (Meridian) $1,250.00 0.4% 5 transfers × $250
Cash/ATM $1,200.00 0.4% 5 ATM withdrawals
Dining $1,195.57 0.4% Restaurants, cafes, bakeries
Transportation $1,005.53 0.4% Gas, Smartrip, Uber, Lyft, parking
Streaming $731.64 0.3% Netflix, Spotify, Disney+
Health/Fitness $528.00 0.2% Peloton
News/Media $300.00 0.1% NYTimes
Software $119.88 0.0% Apple iCloud
Total withdrawals $284,831.76 100.0%
The "Big Four" Structural Outflows
Outflow Annual % of Net Pay ($254,663.36)
Tuition (Cathedral Heights) $120,000.00 47.12%
Mortgage PITI $54,836.04 21.53%
Nanny (net payroll) $46,544.40 18.28%
Credit card autopay $30,630.71 12.03%
Combined Big Four $252,011.15 98.96%
Remaining after Big Four $2,652.21 1.04%

After the Big Four, only $2,652.21 remains from net pay to cover:

  • Utilities: $8,670.83
  • Shopping: $3,067.27
  • Groceries: $1,922.75
  • Dining: $1,195.57
  • Transportation: $1,005.53
  • Subscriptions: $1,679.52
  • Cash/ATM: $1,200.00
  • Total other operating: $18,741.47

Operating deficit: $2,652.21 - $18,741.47 = -$16,089.26

This deficit is partially offset by the fact that some of these categories (dining, shopping, groceries) may flow through the credit card and are already captured in the CapOne autopay. However, the checking account shows BOTH direct debit spending in these categories AND the CapOne autopay, indicating the credit card covers additional spending beyond what's directly debited.

Monthly Cash Flow Table
Month Opening Payroll Deposits Savings In Operating Out CapOne Autopay Savings/Inv Out Net Change Closing
Jan $21,450.00 $20,946.82 $12,000.00 $22,152.00 $2,145.30 $77.90 +$8,571.62 $30,021.62
Feb $30,021.62 $20,946.82 $0.00 $22,600.18 $2,549.65 $250.00 -$4,453.01 $25,568.61
Mar $25,568.61 $20,946.82 $0.00 $21,994.17 $2,288.80 $178.96 -$3,296.71 $22,271.90
Apr $22,271.90 $20,946.82 $0.00 $22,070.79 $2,266.68 $395.65 -$3,536.30 $18,735.60
May $18,735.60 $20,946.82 $15,000.00 $20,666.53 $3,042.67 $1,062.54 +$11,175.08 $29,910.68
Jun $29,910.68 $20,946.82 $0.00 $9,863.61 $2,890.91 $250.00 +$7,942.30 $37,852.98
Jul $37,852.98 $20,946.82 $0.00 $10,060.17 $2,913.85 $168.53 +$7,804.27 $45,657.25
Aug $45,657.25 $20,946.82 $0.00 $22,093.08 $2,279.60 $76.31 -$3,502.17 $42,155.08
Sep $42,155.08 $20,946.82 $0.00 $21,755.01 $2,132.39 $337.15 -$3,197.73 $38,957.35
Oct $38,957.35 $21,238.22 $0.00 $22,415.37 $2,868.63 $211.46 -$4,257.24 $34,700.11
Nov $34,700.11 $22,067.78 $0.00 $21,952.06 $2,301.62 $412.96 -$2,435.89 $32,264.21
Dec $32,264.21 $22,835.98 $0.00 $22,397.98 $2,950.61 $11,403.84 -$13,764.21 $18,500.00

Operating Out = all withdrawals excluding CapOne autopay, savings transfers, and investment transfers.

December includes the $11,251.54 year-end transfer to savings.

Key Monthly Observations

Underwater months (negative net change excluding savings infusions):

  • February: -$4,453.01
  • March: -$3,296.71
  • April: -$3,536.30
  • August: -$3,502.17
  • September: -$3,197.73
  • October: -$4,257.24
  • November: -$2,435.89
  • December: -$13,764.21 (includes $11,251.54 savings transfer)

Surplus months:

  • January: +$8,571.62 (but only because of $12,000 savings transfer)
  • May: +$11,175.08 (but only because of $15,000 savings transfer)
  • June: +$7,942.30 (NO tuition payment — summer break)
  • July: +$7,804.27 (NO tuition payment — summer break)

The June/July tuition gap effect:

Tuition is paid August through May (10 months). June and July have no $12,000 tuition debit. These two months generate approximately $15,746.57 in surplus that partially offsets the deficits in the other 10 months. Without this summer relief, the account would be deeply negative.

Without savings infusions:

  • Total deposits without savings transfers: $254,881.76
  • Total withdrawals without savings transfers out: $273,580.22
  • Net: -$18,698.46
  • Hypothetical ending balance: $21,450.00 - $18,698.46 = $2,751.54

The account would end the year at $2,751.54 instead of $18,500.00 — still technically positive but precariously low.

Savings Transfer Analysis
Metric Amount
Pulled from savings (3 transfers) $27,000.00
Put back to savings (1 transfer, Dec 31) $11,251.54
Net drain from savings $15,748.46
Savings infusions as % of total deposits 9.6%
Months requiring savings infusions January ($12,000), May ($15,000)

The January $12,000 savings transfer appears timed to cover the January tuition payment (both post on the same day, Jan 10). The May $15,000 transfer ($3,000 + $12,000) appears to cover the May tuition payment plus provide a buffer.


SECTION 7: STUDENT LOAN DATA GAP

Source: student_loans

Loan Monthly Payment Annual Total Interest Rate Ending Balance
Niamh (CES-XXXX4471) $1,172.00 $14,064.00 5.05% $141,410.00
Stefan (CES-XXXX8820) $640.00 $7,680.00 4.40% $57,048.00
Combined $1,812.00/mo $21,744.00 $198,458.00

CRITICAL FINDING: Student loan payments ($21,744/year) are NOT visible in the Tidewater checking account.

The 340 transactions contain no debits to Cardinal Education Loan Servicing. This is a material data gap. The loans are being paid from somewhere — possibly:

  • A separate account not provided in the workspace
  • The Capital One credit card (would be embedded in the $30,630.71 autopay)
  • Automatic debit from a different bank account

Impact on true cash flow: If the $21,744 is being paid from a source outside the checking account, the family's true annual cash outflow is at least $21,744 higher than what the checking account shows. If it's being paid via credit card, it's already captured in the CapOne autopay but the credit card balance is being carried (revolving debt).


SECTION 8: SUBSCRIPTION INVENTORY

Source: bank_statement_checking.subscriptions_annual

Subscription Monthly Annual
Netflix $24.99 $299.88
Spotify Family $19.99 $239.88
NYTimes Digital $25.00 $300.00
Apple iCloud $9.99 $119.88
Peloton Membership $44.00 $528.00
Disney+ $15.99 $191.88
Total $139.96/mo $1,679.52

Subscriptions as % of net pay: $1,679.52 / $254,663.36 = 0.66%

Are subscriptions the problem? No. At 0.66% of net pay, subscriptions are immaterial to the family's financial picture. The structural problem is the Big Four outflows consuming 98.96% of net pay.


SECTION 9: AFFORDABILITY ANALYSIS

Can they afford their life as-is?

The short answer: No. They are running a structural deficit, subsidized by savings.

Evidence:

  1. The Big Four consume 98.96% of net pay. After tuition, mortgage, nanny, and credit card payments, only $2,652.21 remains per year ($221/month) for all other living expenses.
  1. Operating expenses far exceed the residual. Utilities alone cost $8,670.83/year. All other operating categories (groceries, shopping, dining, transportation, subscriptions, ATM cash) total $10,070.64. Combined operating need: $18,741.47 against $2,652.21 available = deficit of $16,089.26.
  1. The checking account lost $2,950 even WITH $27,000 in savings infusions. Without those infusions, the account would have ended near $2,751.54 — barely above zero.
  1. The corrected tax bill of $8,288.24 is an unbudgeted cash need. The family prepared their draft return expecting a $1,333.51 refund. Instead, they owe $8,288.24. This is a $9,621.75 swing they have not planned for. The $18,500 year-end balance would drop to $10,211.76 after paying the tax bill.
  1. Student loan payments of $21,744/year are not in the checking account. If these are being paid from a separate account, the true cash drain is even worse than what the checking account shows. If they're on the credit card, the family is revolving $21,744/year in debt.
  1. The family is draining savings. Net $15,748.46 was pulled from savings in 2025. At this rate, savings will be exhausted.
  1. The June/July tuition gap is the only thing keeping the account afloat. Without those two tuition-free months generating ~$15,747 in surplus, the account would be deeply negative even with savings infusions.
Structural Deficit Summary
Component Annual Amount
Net pay deposited $254,663.36
Big Four outflows -$252,011.15
Residual from net pay $2,652.21
Other operating outflows (visible in checking) -$18,741.47
Operating deficit -$16,089.26
Savings infusions (non-recurring) +$27,000.00
Savings returned -$11,251.54
Net savings support +$15,748.46
Net checking account change -$2,950.00
Unbudgeted tax bill -$8,288.24
Student loan payments (not in checking) -$21,744.00
Conclusion

The Okonkwo-Bauers cannot sustain their current lifestyle on their current income. They are:

  • Running a structural operating deficit of ~$16,000/year
  • Subsidizing the deficit with savings withdrawals
  • Facing an unbudgeted $8,288.24 tax bill
  • Carrying $21,744/year in student loan payments from an untraced source
  • Spending 47% of net pay on private school tuition alone

The situation is not a subscription problem or a discretionary spending problem. The Big Four structural commitments (tuition, mortgage, nanny, credit card) leave essentially nothing for everything else. The family is slowly liquidating savings to maintain their lifestyle.


SECTION 10: ASSUMPTIONS AND CAVEATS

  1. Tax year 2025 brackets: Used standard 2025 MFJ brackets with inflation adjustments. The source data does not explicitly state bracket thresholds; the draft 1040 tax calculation was used to verify bracket alignment ($60,836.69 on $321,479.55 taxable income, which cross-checks perfectly).
  1. CTC amount ($4,400): The source states $4,400 for two children. At $2,000 per child, this would be $4,000. The additional $400 may reflect the refundable portion (Additional Child Tax Credit). Used source value as authoritative.
  1. CDCC ($1,200): Used source value. The nanny expenses ($50,400 gross) would support a larger credit ($3,000 for two children × 20% = $600 per child at their income level), but the source states $1,200. Used source value.
  1. DC SUTA ($243.00): Included in fully-loaded nanny cost per source data. This is a state-level employer tax, not a federal tax, and does not appear on Schedule H.
  1. Student loan payment source: The data gap is flagged but cannot be resolved from the provided files. The $21,744 in annual payments must be accounted for in any complete financial picture.
  1. Credit card spending detail: The Capital One statements are not provided. The $30,630.71 annual autopay is a single aggregated outflow. The underlying spending categories (dining, travel, shopping on the card) cannot be analyzed.
  1. Source data internal inconsistency: The critical_calculations section contains pre-computed values that use $152.64 as the full Form 8959 amount instead of $1,421.19. This propagates to incorrect Schedule 2 and total tax figures. The correct values are derived independently above.
  1. Filing deadline: The draft 1040 is marked "NOT YET FILED." The family has time to correct the errors before filing.

APPENDIX: CORRECTED 1040 LINE-BY-LINE REFERENCE

Line Description Corrected Value
1a Wages (Box 1 combined) $368,270.00
2a Tax-exempt interest $1,840.00
2b Taxable interest $3,180.00
3a Qualified dividends $7,860.00
3b Ordinary dividends $9,420.00
7 Capital gain (LTCG) $14,500.00
9 Total income $395,370.00
10 Adjustments $0.00
11 AGI $395,370.00
12 Itemized deductions (Schedule A) $71,390.45
15 Taxable income $323,979.55
16 Tax (QDCGT Worksheet) $61,436.69
19 Child Tax Credit $4,400.00
20 CDCC (Schedule 3) $1,200.00
22 Subtotal after credits $55,836.69
23 Other taxes (Schedule 2) $10,204.19
24 Total tax $66,040.88
25a Box 2 withholding $57,600.00
25c Form 8959 withholding $152.64
25d Total payments $57,752.64
37 Amount you owe $8,288.24

COMMITTED FEDERAL BOTTOM LINE: $8,288.24 BALANCE DUE

Every figure above is traceable to a specific field path in persona_30_task6_persona_files.json. All calculations are independently derived and cross-verified. The source data's critical_calculations section contains pre-computed values that are incorrect (using $152.64 instead of $1,421.19 for Form 8959); the correct values are derived above.

Task 06 · End-to-end software build

DSP Block Emulator

Agentic software engineering · hardware_embedded_and_low_level_systems · dsp_and_signal_hardware

Infer a legacy DSP chip's mask-ROM arithmetic profile from hardware captures and reproduce its exact 16-bit PCM output.

Full brief · 5 source files · 555 lines · 1.0 expert hours

Task BriefThe complete agent-facing instruction+

The mask-programmed legacy fixed-point DSP chip is described in /app/data/DATASHEET.md. Its mask-ROM source was lost, but exact one-step hardware transitions are preserved in /app/data/captures.csv. A recent migration at /app/data/pipeline.py runs without errors and produces valid PCM, but it does not reproduce the physical chip.

Write a self-contained Python 3 program at /app/emulate.py that accepts exactly three command-line arguments and is invoked as:

python3 /app/emulate.py <input.bin> <coeffs.bin> <output.bin>

It must reconstruct the deterministic per-section transition evidenced by every row of /app/data/captures.csv, apply it with the cascade and file contract in /app/data/DATASHEET.md, and write the resulting 16-bit signed little-endian PCM samples to <output.bin>. Use only the Python standard library.

Use /app/data/pipeline.py as the runnable starting point. Create /app/emulate.py and /app/output.bin before detailed capture analysis, keep both artifacts working while refining the transition, and leave the best working reconstruction in place even if some captures remain unresolved.

The coefficient file begins with a little-endian uint16 section count, followed by that many sections. Each section contains five little-endian int16 Q15 coefficients in the order b0, b1, b2, a1, a2.

Run your program on the shipped data to create /app/output.bin:

python3 /app/emulate.py /app/data/input.bin /app/data/coeffs.bin /app/output.bin

Do not modify any files in /app/data/. Do not read /tests.

The completed /app/output.bin must satisfy these numbered criteria:

  1. It exists and contains the same number of 16-bit signed little-endian samples as /app/data/input.bin.
  2. Every sample is the exact output the legacy chip would produce for /app/data/input.bin and /app/data/coeffs.bin under the behavior evidenced by /app/data/captures.csv and the cascade contract in /app/data/DATASHEET.md.
  3. /app/emulate.py uses only the Python standard library.
Implementation & Verification Code5 authored source files: solution, public tests, and build environment+
Source Browsersolution/solve.py
5 files · 555 lines
"""Oracle for the legacy mask-programmed fixed-point DSP chip."""
import struct
import sys

Q = 15
HALF = 1 << (Q - 1)
MASK = (1 << Q) - 1
STATE_INIT = [(12345, -6789), (11111, -22222)]


def mul_cell(coefficient, value):
    product = coefficient * value
    return product - (product & 0xF)  # M1


def round_cell(value):
    quotient = value >> Q
    remainder = value & MASK
    if remainder < HALF:
        return quotient
    if remainder > HALF:
        return quotient + 1
    return quotient if quotient & 1 else quotient + 1  # Q1


def rail_cell(value):
    if value > 32767:
        return 32767
    if value < -32767:
        return -32767
    return value  # L1


def feedback_cell(rounded, emitted):
    return emitted  # F0


def route_cell(u1, u2, residual):
    return u1 + residual, u2 - residual  # E1


def state_cell(value):
    if value > 536870911:
        return 536870911
    if value < -536870912:
        return -536870912
    return value  # W1


def run_section(x, coefficients, d1, d2):
    b0, b1, b2, a1, a2 = coefficients
    acc = mul_cell(b0, x) + d1
    rounded = round_cell(acc)
    emitted = rail_cell(rounded)
    feedback = feedback_cell(rounded, emitted)
    residual = (emitted << Q) - acc

    u1 = mul_cell(b1, x) - mul_cell(a1, feedback) + d2
    u2 = mul_cell(b2, x) - mul_cell(a2, feedback)
    u1, u2 = route_cell(u1, u2, residual)
    return emitted, state_cell(u1), state_cell(u2)


def emulate(samples, coefficient_sections):
    states = list(STATE_INIT[:len(coefficient_sections)])
    output = []
    for sample in samples:
        emitted = sample
        for index, coefficients in enumerate(coefficient_sections):
            emitted, d1, d2 = run_section(
                emitted,
                coefficients,
                states[index][0],
                states[index][1],
            )
            states[index] = (d1, d2)
        output.append(emitted)
    return output


def read_samples(path):
    with open(path, "rb") as stream:
        data = stream.read()
    if len(data) % 2:
        raise ValueError("input PCM has an incomplete sample")
    return struct.unpack(f"<{len(data) // 2}h", data)


def read_coefficients(path):
    with open(path, "rb") as stream:
        data = stream.read()
    if len(data) < 2:
        raise ValueError("coefficient file too short")
    section_count = struct.unpack_from("<H", data, 0)[0]
    if section_count not in (1, 2):
        raise ValueError("section count must be one or two")
    if len(data) != 2 + section_count * 10:
        raise ValueError("coefficient file has an invalid size")
    return [
        struct.unpack_from("<5h", data, 2 + index * 10)
        for index in range(section_count)
    ]


def write_samples(path, samples):
    with open(path, "wb") as stream:
        stream.write(struct.pack(f"<{len(samples)}h", *samples))


if __name__ == "__main__":
    if len(sys.argv) != 4:
        print("Usage: python3 solve.py <input.bin> <coeffs.bin> <output.bin>", file=sys.stderr)
        sys.exit(1)
    write_samples(
        sys.argv[3],
        emulate(read_samples(sys.argv[1]), read_coefficients(sys.argv[2])),
    )
Evaluation ContractCategory, tested runtime, effort, and required artifacts+
Public identifierevaluation/dsp-block-emulator
Categoryhardware_embedded_and_low_level_systems
Subcategorydsp_and_signal_hardware
Expert estimate1.0 hours
Model testedGPT-5.4
Agent testedTerminus-2
Required artifacts/app/output.bin, /app/emulate.py
Difficulty DesignWhy the benchmark discriminates between plausible and correct work+

This task deliberately combines three the technical evaluation program stumping patterns. Pattern D, evidence-forced reverse engineering: the agent-visible notes preserve the architecture but not the proprietary mask logic, so the exact transition must be recovered from 35 physical one-step captures. Pattern B, wrong-default lure: /app/data/pipeline.py is a clean, plausible DF2T migration that produces normal PCM and is correct on some conventions, encouraging a solver to patch the first visible mismatch and stop. Pattern G, many independent bugs under an all-or-nothing artifact: the physical chip independently clears multiplier guard bits, uses a custom tie rule, reserves one output code, has a specific overload-feedback choice, routes quantization residual into both state paths, and clips narrower state registers. Missing any one rule corrupts a hidden stream and the exact artifact fails. The runnable migration makes an early best-effort artifact feasible within the one-hour evaluation, while exact capture reconstruction remains the crux. Silent failure and no-self-check amplify the trap because every near-miss remains valid PCM and public output ground truth is unavailable. The capture contains repeated isolating evidence for every convention, so a firmware expert using direct differential analysis can derive the result without guessing. All PCM, coefficients, transitions, and fixtures are deterministic synthetic data modeling logic-analyzer-driven compatibility work on an undocumented mask-ROM revision.

Reference ApproachThe intended solution strategy+

Compare pipeline.py's predicted one-step transitions against captures.csv rather than trusting textbook DF2T defaults. The residual patterns reveal that every coefficient product clears its low four bits toward negative infinity; exact Q15 ties are forced odd; the negative output rail is -32767; the emitted rail value itself is fed back; the quantization residual is added to the first state path and subtracted from the second; and state registers saturate to signed 30-bit bounds. Implement those recovered rules with the documented boot states, two's-complement coefficient decoding, cascade order, and persistent d1/d2 state, then run the emulator on the shipped files to produce /app/output.bin.

Verification DesignHow the produced artifacts are independently checked+

The solving agent runs as an unprivileged account that can create the two required /app artifacts but cannot modify /app/data or the verifier toolchain. The verifier requires regular, non-symlink /app/emulate.py and /app/output.bin files, validates the public artifact shape, proves its protected transition matches every capture row, enforces the standard-library-only and exact-three-argument contracts, re-runs the script, and requires exact verifier-side 16-bit PCM. Every candidate re-run drops again to a separate unprivileged uid with /tests root-only, isolated Python startup, a runtime import guard, and bounded subprocess timeouts; reference PCM is computed only in the protected parent process and is never stored at a candidate-readable path. Exact public and hidden comparisons necessarily reject output copied from the migration. Five protected fixed streams isolate product clearing and residual routing, the reserved negative rail, exact tie behavior, narrow-state clipping followed by recovery, and overload feedback. A further fixed in-memory case has no pre-baked expected-output file. Author-side mutation analysis confirms that replacing any one recovered convention with the corresponding plausible migration behavior changes at least one graded stream. Exact comparison is justified because the physical transition evidence and all integer formats are exact.

Task 07 · End-to-end software build

Veiled Current

Agentic software engineering · Games Puzzles and Interactive Simulation · Game AI and Strategy

Implement a finite-horizon adversarial grid solver with a persistent hidden current and public landing observations.

Full brief · 6 source files · 571 lines · 4.0 expert hours

Task BriefThe complete agent-facing instruction+

Veiled Current is a two-player zero-sum stochastic game on a rectangular grid. Each cell is one of:

  • . — floor
  • # — wall
  • T — target
  • G — trap

At the start of each position, Nature samples one hidden current, L or R. The sampled current remains fixed for the entire game. Both players know the prior probability of L, but neither player ever observes the current itself.

The token starts on the specified floor cell. The players alternate moving it, with player 1 moving first. On a turn, the mover publicly chooses north, south, west, or east. A player may randomize this choice. Movement is then sampled relative to the chosen direction:

| Hidden current | Intended direction | Left turn | Right turn | |---|---:|---:|---:| | L | 50% | 45% | 5% | | R | 50% | 5% | 45% |

For example, relative to north, west is left and east is right. A sampled move that would leave the grid or enter # leaves the token in its current cell.

After every move, both players observe the chosen direction and the token's resulting cell. They do not observe whether the intended, left, or right branch was sampled. If multiple branches lead to the same resulting cell, that single cell is the entire observation. Both players remember the full public history and may base later choices on it, but neither may condition a choice on the hidden current or an unobserved movement branch.

The game ends immediately upon entering T or G. Entering T is a player-1 win; entering G is a player-1 loss. If neither is reached within the stated number of moves, player 1 loses. Player 1 maximizes the win probability and player 2 minimizes it.

/app/data/positions.txt contains one position per line:

rows,cols,horizon,prior_L_numerator,prior_L_denominator,start_row,start_col:grid_string

Coordinates are zero-based. grid_string is a row-major concatenation of exactly rows * cols characters. The prior probability of current L is prior_L_numerator / prior_L_denominator. All graded grids are at most 4-by-5, and all horizons are at most 8.

/app/data/expected.txt contains the correct player-1 win probability for each visible position in the same order.

Write a self-contained solver at /app/solve.py. Running python3 /app/solve.py with no arguments or stdin must read /app/data/positions.txt and write one probability per line to /app/output.txt, with at least six digits after the decimal point.

Graded positions follow the same rules. Outputs are accepted within absolute tolerance 1e-6. Compute from the rules rather than hard-coding positions. Do not modify /app/data/, read or import /tests, use third-party packages, or make network requests. Each solver invocation must finish within 120 seconds.

Implementation & Verification Code6 authored source files: solution, public tests, and build environment+
Source Browsersolution/solve.py
6 files · 571 lines
"""Golden solver for the Veiled Current finite-horizon game."""

from fractions import Fraction
from functools import lru_cache


DIRECTIONS = ((-1, 0), (1, 0), (0, -1), (0, 1))
PERPENDICULAR = (
    ((0, -1), (0, 1)),
    ((0, 1), (0, -1)),
    ((1, 0), (-1, 0)),
    ((-1, 0), (1, 0)),
)
MODE_WEIGHTS = (
    (Fraction(1, 2), Fraction(9, 20), Fraction(1, 20)),
    (Fraction(1, 2), Fraction(1, 20), Fraction(9, 20)),
)


def parse_positions(path):
    positions = []
    with open(path, "r", encoding="utf-8") as handle:
        for raw_line in handle:
            line = raw_line.strip()
            if not line:
                continue
            metadata, grid = line.split(":", 1)
            values = tuple(map(int, metadata.split(",")))
            if len(values) != 7:
                raise ValueError("position metadata must contain seven integers")
            rows, cols, horizon, numerator, denominator, start_r, start_c = values
            if len(grid) != rows * cols:
                raise ValueError("grid length does not match dimensions")
            positions.append(
                (rows, cols, horizon, Fraction(numerator, denominator),
                 start_r, start_c, grid)
            )
    return positions


def solve_position(rows, cols, horizon, prior, start_r, start_c, grid):
    def land(cell, delta):
        row, col = divmod(cell, cols)
        new_r, new_c = row + delta[0], col + delta[1]
        if not (0 <= new_r < rows and 0 <= new_c < cols):
            return cell
        destination = new_r * cols + new_c
        return cell if grid[destination] == "#" else destination

    transitions = {}
    for cell, symbol in enumerate(grid):
        if symbol in "#TG":
            continue
        for action, forward in enumerate(DIRECTIONS):
            left, right = PERPENDICULAR[action]
            destinations = (land(cell, forward), land(cell, left), land(cell, right))
            for mode in (0, 1):
                distribution = {}
                for destination, probability in zip(
                    destinations, MODE_WEIGHTS[mode]
                ):
                    distribution[destination] = (
                        distribution.get(destination, Fraction(0)) + probability
                    )
                transitions[cell, action, mode] = distribution

    @lru_cache(maxsize=None)
    def value(cell, player, remaining, belief):
        if grid[cell] == "T":
            return Fraction(1)
        if grid[cell] == "G" or remaining == 0:
            return Fraction(0)

        action_values = []
        for action in range(4):
            under_l = transitions[cell, action, 0]
            under_r = transitions[cell, action, 1]
            expectation = Fraction(0)
            for destination in under_l.keys() | under_r.keys():
                likelihood_l = under_l.get(destination, Fraction(0))
                likelihood_r = under_r.get(destination, Fraction(0))
                observation_probability = (
                    belief * likelihood_l + (1 - belief) * likelihood_r
                )
                if observation_probability:
                    posterior = (
                        belief * likelihood_l / observation_probability
                    )
                    expectation += observation_probability * value(
                        destination, 1 - player, remaining - 1, posterior
                    )
            action_values.append(expectation)

        return max(action_values) if player == 0 else min(action_values)

    start = start_r * cols + start_c
    return value(start, 0, horizon, prior)


def main():
    positions = parse_positions("/app/data/positions.txt")
    answers = [solve_position(*position) for position in positions]
    with open("/app/output.txt", "w", encoding="utf-8") as handle:
        handle.write("\n".join(f"{float(answer):.10f}" for answer in answers) + "\n")


if __name__ == "__main__":
    main()
Evaluation ContractCategory, tested runtime, effort, and required artifacts+
Public identifierevaluation/veiled-current
CategoryGames Puzzles and Interactive Simulation
SubcategoryGame AI and Strategy
Expert estimate4.0 hours
Model testedGPT-5.4
Agent testedTerminus-2
Required artifacts/app/solve.py, /app/output.txt
Difficulty DesignWhy the benchmark discriminates between plausible and correct work+

The state visible on the board is not Markov by itself. Nature samples one current that persists for the full game, actions reveal information only through the landing cell, and walls or boundaries can merge intended and drift branches into the same public observation. A solver must therefore carry a common posterior through an alternating max/min finite-horizon game and update it using the total likelihood of each observed landing cell. Three attractive shortcuts all run quickly and look locally reasonable: averaging the two currents once and freezing the prior, treating the sampled movement branch as public, or solving two fully observed games and averaging their optimized values. Compact hidden fixtures mutation-test all three with clear numeric margins. This mirrors work by researchers and verification or security engineers who synthesize policies for imperfect-information stochastic games, partially observed controllers, and security patrol planning. Every position is an authored synthetic grid selected by mutation testing to isolate belief-update errors while remaining representative of those information-state decisions. Grids are at most 4-by-5 and horizons at most eight, so the challenge is the coupled information structure rather than state volume, iteration count, or timeout pressure.

Reference ApproachThe intended solution strategy+

Represent the public information state by token cell, side to move, remaining horizon, and the rational posterior probability of current L. For each action, aggregate intended/left/right probabilities by landing cell separately under L and R. The predictive probability of a public landing is the posterior-weighted sum of those two likelihoods. Recurse from that landing with one fewer move and update the posterior by Bayes' rule using the aggregated landing likelihood. Player 1 selects the maximum action value and player 2 the minimum. Terminal target and trap values are one and zero, and horizon zero is zero. Exact fractions make equivalent public histories share memoized states and avoid numerical drift.

Verification DesignHow the produced artifacts are independently checked+

The verifier checks artifact existence, standard-library-only code, source and runtime isolation from /tests, visible outputs, hidden-fixture integrity, and hidden outputs against an independent exact-rational oracle. A public-observation regression confirms that colliding branches are aggregated before belief updating. Mutation calibration requires every hidden fixture to reject the frozen-prior model, the branch-revealing model, and the fully-observed-current model by more than the disclosed tolerance. The submitted solver is executed as uid/gid 65534 with no_new_privs, closed descriptors, root-only verifier files, and a 120-second per-invocation cap. The full verifier has a 300-second budget; expected execution is far below it.

Task 08 · End-to-end software build

Blackout Cascade

Agentic software engineering · Games Puzzles and Interactive Simulation · Interactive text games

Generate an optimal-play replay certificate for a control-fracturing, multi-table variant of the partizan game Blackout.

Full brief · 5 source files · 1,146 lines · 4.0 expert hours

Task BriefThe complete agent-facing instruction+

Produce an optimal-play replay certificate for the Blackout Cascade relay in /app/data/positions.txt. The certificate is used to compare deterministic rules-engine implementations.

Input. /app/data/positions.txt is whitespace-separated. Each table is p <n_lights> and base light bits; a <n_alloff> and AllOff rows; o <n_oneon> and OneOn rows; then ids <n_ids> and the row-ordered AllOff then OneOn ids. The a rows and first n_alloff ids form AllOff's switch pool; the o rows and remaining ids form OneOn's pool. A mover may select or nominate only an available switch from its own pool. Ids are unique within a table. Rows have <n_lights> bits; 1 toggles that bulb when on. Off consumes without toggling. Bit lists put bulb 0 first; their integer mask is $\sum_i bit_i 2^i$.

Per-table rules. Table 0 starts with AllOff; later tables start with the opponent of the player whose turn ended the prior table. Each table begins intact. Before selecting, the mover may irreversibly fracture that table. Intact: the mover selects a switch and off/on. Fractured: the mover nominates exactly two available switches, or the sole remaining switch; the opponent selects a nominee and off/on. The selected switch is consumed and the other remains. Fracture and nomination consume no move.

A relay-wide blackout run starts at 0, increments when a consumed switch leaves mask 0, and resets after a nonzero result; initial darkness does not count. At run 2, AllOff wins immediately and the run resets before the next table. Earlier endings carry run 0 or 1 forward. Otherwise, AllOff with no switch ends the table, winning exactly at mask 0. OneOn with no switch loses at mask 0 and otherwise passes.

Deterministic choice. Each chooser prefers its own table win, minimizing remaining moves when winning and maximizing them when losing. Exact ties prefer intact over fracture, then lowest row index and off before on for direct selections and responses, or the lexicographically lowest tuple of nominated row indices. Only consumption or pass counts as a move.

Relay entanglement. A switch used on table $i$ is unavailable on tables $i+1$ and $i+2$, then available again. An id keeps the same owner everywhere it appears. Table 0 uses its given lights; table $i>0$ starts from given_lights(i) XOR final_lights(i-1), restricted to its current n_lights low bits. Let used_alloff count AllOff switches played on a table. Let unused_oneon count OneOn switches available at its start but not played there. Only the table winner receives points: if AllOff wins, points_awarded is the single integer n_lights + unused_oneon; if OneOn wins, it is popcount(final_lights) + used_alloff. The relay winner has the higher cumulative score; a tie goes to the last table winner.

Output. /app/output.json must be one JSON object with exactly:

  • relay_winner: "AllOff" or "OneOn".
  • scores: exactly {"AllOff": <int>, "OneOn": <int>}.
  • tables: one entry per input table. Each entry has exactly initial_lights (effective integer mask), starting_player, starting_blackout, winner, final_lights, blackout_carry, next_player, points_awarded (the winner-only integer defined above), used_alloff, unused_oneon, and line.

Each line entry has exactly mover, mode, nominees, selected, action, and lights_after. For an intact move, mode is "intact", nominees is [], selected is the switch id, and action is "off" or "on". For a fractured move, mode is "fractured", nominees lists the nominated ids in row order, and selected is the opponent's chosen nominee. For a pass, use "pass", [], null, and "pass" respectively. Player fields use "AllOff" or "OneOn"; masks and counts are integers.

Create /app/solve.py; running python3 /app/solve.py with no arguments or stdin must read /app/data/positions.txt and write /app/output.json. Use only Python's standard library. Do not modify /app/data/, access /tests, use the network, or hard-code answers. Fresh hidden relays follow these rules.

Implementation & Verification Code5 authored source files: solution, public tests, and build environment+
Source Browsersolution/solve.py
5 files · 1,146 lines
#!/usr/bin/env python3
"""Oracle solver for Blackout Cascade.

Reads /app/data/positions.txt and writes an optimal-play replay certificate to
/app/output.json.
"""
from __future__ import annotations

import json
from functools import lru_cache
from itertools import combinations
from pathlib import Path

DATA = Path("/app/data/positions.txt")
OUT = Path("/app/output.json")

ALLOFF_TURN = 0
ONEON_TURN = 1
ALLOFF_WINNER = "AllOff"
ONEON_WINNER = "OneOn"

BLACKOUT_CONFIRM_MOVES = 2
BLACKOUT_CARRIES_ACROSS_TABLES = True
STARTER_CARRIES_ACROSS_TABLES = True
FATIGUE_TABLES = 2
XOR_INHERITANCE = True
MASK_INHERITED_LIGHTS = True
ALLOFF_RESOURCE_SCORING = True
ONEON_RESOURCE_SCORING = True

# Local fixture generation mutates these switches to prove that the hidden bank
# distinguishes the fracture rules.  Canonical submissions implement the values
# below directly; the names are not part of the task interface.
FRACTURE_AVAILABLE = True
FRACTURE_FORCED = False
NOMINATION_SIZE = 2
RESPONDER_USES_MOVER_PREFERENCE = False
FRACTURE_COUNTS_AS_MOVE = False
INTACT_TIE_PREFERRED = True
HIGH_NOMINATION = False
USE_OUTCOME_DISTANCE = True
ALLOFF_HIGH_INDEX = False
ALLOFF_ON_FIRST = False
ONEON_HIGH_INDEX = False
ONEON_ON_FIRST = False


def parse_relay(path: Path):
    """Yield per-table (p, lights, alloff_masks, oneon_masks, ids)."""
    text = path.read_text(encoding="utf-8")
    tokens = text.split()
    it = iter(tokens)

    def read_bits(width: int) -> int:
        mask = 0
        for i in range(width):
            if int(next(it)):
                mask |= 1 << i
        return mask

    while True:
        try:
            header = next(it)
        except StopIteration:
            return
        if header != "p":
            raise ValueError(f"expected table header 'p', got {header!r}")
        p = int(next(it))
        lights = read_bits(p)
        if next(it) != "a":
            raise ValueError("expected 'a' after lights")
        a = int(next(it))
        alloff = [read_bits(p) for _ in range(a)]
        if next(it) != "o":
            raise ValueError("expected 'o' after AllOff switches")
        o = int(next(it))
        oneon = [read_bits(p) for _ in range(o)]
        if next(it) != "ids":
            raise ValueError("expected 'ids' after OneOn switches")
        n_ids = int(next(it))
        ids = [next(it) for _ in range(n_ids)]
        if len(ids) != a + o:
            raise ValueError(f"ids count {len(ids)} != a+o {a+o}")
        yield p, lights, alloff, oneon, ids


# Module-level switch tables used by the memoised state_value.
_alloff_masks: tuple[int, ...] = ()
_oneon_masks: tuple[int, ...] = ()


def outcome_distance_key(turn: int, result: tuple[bool, int]):
    """Order a result by the selecting player's outcome and distance."""
    alloff_wins, distance = result
    current_wins = alloff_wins if turn == ALLOFF_TURN else not alloff_wins
    return (
        0 if current_wins else 1,
        (
            distance if current_wins else -distance
        ) if USE_OUTCOME_DISTANCE else 0,
    )


def selection_key(turn: int, idx: int, action: str):
    """Order a direct selection by the chooser's deterministic tie-break."""
    high_index = (
        ALLOFF_HIGH_INDEX if turn == ALLOFF_TURN else ONEON_HIGH_INDEX
    )
    on_first = ALLOFF_ON_FIRST if turn == ALLOFF_TURN else ONEON_ON_FIRST
    return (
        -idx if high_index else idx,
        0 if (action == "on") == on_first else 1,
    )


def preference_key(
    turn: int,
    result: tuple[bool, int],
    idx: int = 0,
    action: str = "off",
):
    """Order outcome, distance, switch index, then off before on."""
    return outcome_distance_key(turn, result) + selection_key(
        turn, idx, action
    )


def available_indices(mask: int) -> tuple[int, ...]:
    """Return the set-bit indices in ascending order."""
    result = []
    while mask:
        bit = mask & -mask
        result.append(bit.bit_length() - 1)
        mask ^= bit
    return tuple(result)


def nomination_sets(mask: int):
    """Return the legal one- or two-switch nominations."""
    indices = available_indices(mask)
    width = min(NOMINATION_SIZE, len(indices))
    return combinations(indices, width)


def nomination_key(nomination: tuple[int, ...]):
    """Order a nomination by the disclosed row-index tie-break."""
    return (
        tuple(-idx for idx in nomination)
        if HIGH_NOMINATION
        else nomination
    )


def played_child(
    lights: int,
    a_mask: int,
    o_mask: int,
    turn: int,
    blackout_run: int,
    fractured: bool,
    idx: int,
    action: str,
):
    """Evaluate one consumed switch and return its child result and state."""
    if turn == ALLOFF_TURN:
        switch_mask = _alloff_masks[idx]
        next_a_mask = a_mask ^ (1 << idx)
        next_o_mask = o_mask
    else:
        switch_mask = _oneon_masks[idx]
        next_a_mask = a_mask
        next_o_mask = o_mask ^ (1 << idx)

    new_lights = lights if action == "off" else lights ^ switch_mask
    new_run = blackout_run + 1 if new_lights == 0 else 0
    if new_run >= BLACKOUT_CONFIRM_MOVES:
        result = (True, 0)
    else:
        result = state_result(
            new_lights,
            next_a_mask,
            next_o_mask,
            ONEON_TURN if turn == ALLOFF_TURN else ALLOFF_TURN,
            new_run,
            fractured,
        )
    return result, new_lights, new_run


def nomination_response(
    lights: int,
    a_mask: int,
    o_mask: int,
    turn: int,
    blackout_run: int,
    nomination: tuple[int, ...],
):
    """Return the opponent's chosen result, switch, action, and new state."""
    response_turn = (
        turn
        if RESPONDER_USES_MOVER_PREFERENCE
        else ONEON_TURN if turn == ALLOFF_TURN else ALLOFF_TURN
    )
    choices = []
    for idx in nomination:
        for action in ("off", "on"):
            result, new_lights, new_run = played_child(
                lights,
                a_mask,
                o_mask,
                turn,
                blackout_run,
                True,
                idx,
                action,
            )
            choices.append((result, idx, action, new_lights, new_run))
    return min(
        choices,
        key=lambda choice: preference_key(
            response_turn, choice[0], choice[1], choice[2]
        ),
    )


@lru_cache(maxsize=None)
def state_result(
    lights: int,
    a_mask: int,
    o_mask: int,
    turn: int,
    blackout_run: int = 0,
    fractured: bool = False,
) -> tuple[bool, int]:
    """Return (AllOff wins, consumed switches/passes to termination)."""
    if turn == ALLOFF_TURN:
        if a_mask == 0:
            return lights == 0, 0
    else:
        if o_mask == 0:
            if lights == 0:
                return True, 0
            winner, distance = state_result(
                lights, a_mask, o_mask, ALLOFF_TURN, 0, fractured
            )
            return winner, distance + 1

    own_mask = a_mask if turn == ALLOFF_TURN else o_mask
    choices = []

    # Before fracture, the mover may retain ordinary control.  This branch is
    # absent in the "always fracture immediately" local mutation.
    if not fractured and not FRACTURE_FORCED:
        for idx in available_indices(own_mask):
            for action in ("off", "on"):
                result, _, _ = played_child(
                    lights,
                    a_mask,
                    o_mask,
                    turn,
                    blackout_run,
                    False,
                    idx,
                    action,
                )
                key = (
                    outcome_distance_key(turn, result)
                    + (0 if INTACT_TIE_PREFERRED else 1,)
                    + selection_key(turn, idx, action)
                )
                choices.append((key, result))

    # A fracture is available on any unfractured turn; after fracture,
    # nomination is the only move form.
    if fractured or FRACTURE_AVAILABLE:
        for nomination in nomination_sets(own_mask):
            result, _, _, _, _ = nomination_response(
                lights,
                a_mask,
                o_mask,
                turn,
                blackout_run,
                nomination,
            )
            if not fractured and FRACTURE_COUNTS_AS_MOVE:
                result = (result[0], result[1] + 1)
            mode_rank = (
                ()
                if fractured
                else (1 if INTACT_TIE_PREFERRED else 0,)
            )
            key = (
                outcome_distance_key(turn, result)
                + mode_rank
                + nomination_key(nomination)
            )
            choices.append((key, result))

    _, (winner, distance) = min(choices, key=lambda choice: choice[0])
    return winner, distance + 1


def play_deterministic(
    lights: int,
    a_mask: int,
    o_mask: int,
    turn: int,
    a_count: int,
    ids: list[str],
    starting_blackout_run: int = 0,
):
    """Play a single table under deterministic optimal-play tie-breaks.

    Returns (winner, used_a_mask, used_o_mask, final_lights, moves,
    blackout_run, next_starter).
    moves is a list of internal replay-event dictionaries.
    ids[a_count + j] is the ID of OneOn local switch j.
    """
    used_a = 0
    used_o = 0
    moves = []
    blackout_run = starting_blackout_run
    fractured = False

    while True:
        if turn == ALLOFF_TURN:
            if a_mask == 0:
                winner = ALLOFF_WINNER if lights == 0 else ONEON_WINNER
                return (
                    winner,
                    used_a,
                    used_o,
                    lights,
                    moves,
                    blackout_run,
                    ONEON_TURN,
                )

        else:  # ONEON_TURN
            if o_mask == 0:
                if lights == 0:
                    return (
                        ALLOFF_WINNER,
                        used_a,
                        used_o,
                        lights,
                        moves,
                        blackout_run,
                        ALLOFF_TURN,
                    )
                moves.append({
                    "player": ONEON_WINNER,
                    "switch_id": None,
                    "action": "pass",
                    "lights_after": lights,
                    "fractured": fractured,
                    "nominees": (),
                })
                blackout_run = 0
                turn = ALLOFF_TURN
                continue

        own_mask = a_mask if turn == ALLOFF_TURN else o_mask
        choices = []

        if not fractured and not FRACTURE_FORCED:
            for idx in available_indices(own_mask):
                for action in ("off", "on"):
                    result, new_lights, new_run = played_child(
                        lights,
                        a_mask,
                        o_mask,
                        turn,
                        blackout_run,
                        False,
                        idx,
                        action,
                    )
                    key = (
                        outcome_distance_key(turn, result)
                        + (0 if INTACT_TIE_PREFERRED else 1,)
                        + selection_key(turn, idx, action)
                    )
                    choices.append(
                        (
                            key,
                            False,
                            (),
                            idx,
                            action,
                            new_lights,
                            new_run,
                        )
                    )

        if fractured or FRACTURE_AVAILABLE:
            for nomination in nomination_sets(own_mask):
                response = nomination_response(
                    lights,
                    a_mask,
                    o_mask,
                    turn,
                    blackout_run,
                    nomination,
                )
                result, idx, action, new_lights, new_run = response
                keyed_result = result
                if not fractured and FRACTURE_COUNTS_AS_MOVE:
                    keyed_result = (result[0], result[1] + 1)
                mode_rank = (
                    ()
                    if fractured
                    else (1 if INTACT_TIE_PREFERRED else 0,)
                )
                key = (
                    outcome_distance_key(turn, keyed_result)
                    + mode_rank
                    + nomination_key(nomination)
                )
                choices.append(
                    (
                        key,
                        True,
                        nomination,
                        idx,
                        action,
                        new_lights,
                        new_run,
                    )
                )

        (
            _,
            chosen_fractured,
            nomination,
            idx,
            action,
            lights,
            blackout_run,
        ) = min(
            choices, key=lambda choice: choice[0]
        )
        fractured = fractured or chosen_fractured

        bit = 1 << idx
        if turn == ALLOFF_TURN:
            used_a |= bit
            a_mask ^= bit
            player = ALLOFF_WINNER
            switch_id = ids[idx]
            next_turn = ONEON_TURN
        else:
            used_o |= bit
            o_mask ^= bit
            player = ONEON_WINNER
            switch_id = ids[a_count + idx]
            next_turn = ALLOFF_TURN

        moves.append({
            "player": player,
            "switch_id": switch_id,
            "action": action,
            "lights_after": lights,
            "fractured": fractured,
            "nominees": tuple(
                ids[n] if turn == ALLOFF_TURN else ids[a_count + n]
                for n in nomination
            ),
        })
        if blackout_run >= BLACKOUT_CONFIRM_MOVES:
            return (
                ALLOFF_WINNER,
                used_a,
                used_o,
                lights,
                moves,
                blackout_run,
                next_turn,
            )
        turn = next_turn


def player_name(turn: int) -> str:
    """Return the public name for a player code."""
    return ALLOFF_WINNER if turn == ALLOFF_TURN else ONEON_WINNER


def public_move(move: dict) -> dict:
    """Convert one internal move into the replay-certificate schema."""
    action = move["action"]
    selected = move["switch_id"]
    if action == "pass":
        mode = "pass"
        nominees = []
    elif move.get("fractured", False):
        mode = "fractured"
        nominees = list(move.get("nominees", ()))
    else:
        mode = "intact"
        nominees = []
    return {
        "mover": move["player"],
        "mode": mode,
        "nominees": nominees,
        "selected": selected,
        "action": action,
        "lights_after": move["lights_after"],
    }


# Exposed for mutation-aware fixture generation only; not part of output schema.
_last_relay_totals: tuple[int, int] = (0, 0)
_last_last_winner: str = ""


def compute_outputs(tables: list):
    """Solve the full cascade and return its replay certificate."""
    global _last_relay_totals, _last_last_winner
    a_total = 0
    o_total = 0
    table_reports = []

    unavailable_until: dict[str, int] = {}
    blackout_run = 0
    starter = ALLOFF_TURN

    for tidx, (p, base_lights, alloff, oneon, ids) in enumerate(tables):
        a_count = len(alloff)
        o_count = len(oneon)

        global _alloff_masks, _oneon_masks
        _alloff_masks = tuple(alloff)
        _oneon_masks = tuple(oneon)
        state_result.cache_clear()

        a_mask = 0
        for j in range(a_count):
            sid = ids[j]
            if tidx >= unavailable_until.get(sid, 0):
                a_mask |= 1 << j
        o_mask = 0
        for j in range(o_count):
            sid = ids[a_count + j]
            if tidx >= unavailable_until.get(sid, 0):
                o_mask |= 1 << j

        if tidx == 0:
            lights = base_lights
        else:
            inherited = prev_final_lights if XOR_INHERITANCE else 0
            lights = base_lights ^ inherited
            if MASK_INHERITED_LIGHTS:
                lights &= (1 << p) - 1
        initial_lights = lights
        starting_player = player_name(starter)
        starting_blackout = blackout_run

        (
            winner,
            used_a,
            used_o,
            final_lights,
            moves,
            terminal_blackout_run,
            next_starter,
        ) = play_deterministic(
                lights,
                a_mask,
                o_mask,
                starter,
                a_count,
                ids,
                blackout_run,
            )
        blackout_carry = (
            terminal_blackout_run
            if BLACKOUT_CARRIES_ACROSS_TABLES
            and terminal_blackout_run < BLACKOUT_CONFIRM_MOVES
            else 0
        )
        next_starter = (
            next_starter
            if STARTER_CARRIES_ACROSS_TABLES
            else ALLOFF_TURN
        )

        for j in range(a_count):
            if used_a & (1 << j):
                unavailable_until[ids[j]] = tidx + FATIGUE_TABLES + 1
        for j in range(o_count):
            if used_o & (1 << j):
                unavailable_until[ids[a_count + j]] = (
                    tidx + FATIGUE_TABLES + 1
                )

        used_alloff_count = used_a.bit_count()
        unused_oneon_count = (o_mask & ~used_o).bit_count()
        # Scoring rewards board control and denied opposing resources.
        if winner == ALLOFF_WINNER:
            points = p + (
                unused_oneon_count if ALLOFF_RESOURCE_SCORING else 0
            )
            a_total += points
        else:
            points = final_lights.bit_count() + (
                used_alloff_count if ONEON_RESOURCE_SCORING else 0
            )
            o_total += points

        table_reports.append({
            "initial_lights": initial_lights,
            "starting_player": starting_player,
            "starting_blackout": starting_blackout,
            "winner": winner,
            "final_lights": final_lights,
            "blackout_carry": blackout_carry,
            "next_player": player_name(next_starter),
            "points_awarded": points,
            "used_alloff": used_alloff_count,
            "unused_oneon": unused_oneon_count,
            "line": [public_move(move) for move in moves],
        })

        prev_final_lights = final_lights
        blackout_run = blackout_carry
        starter = next_starter

    if a_total > o_total:
        relay_winner = ALLOFF_WINNER
    elif o_total > a_total:
        relay_winner = ONEON_WINNER
    else:
        relay_winner = table_reports[-1]["winner"]

    _last_relay_totals = (a_total, o_total)
    _last_last_winner = table_reports[-1]["winner"]

    return {
        "relay_winner": relay_winner,
        "scores": {
            ALLOFF_WINNER: a_total,
            ONEON_WINNER: o_total,
        },
        "tables": table_reports,
    }


def main() -> None:
    tables = list(parse_relay(DATA))
    result = compute_outputs(tables)
    OUT.write_text(json.dumps(result, indent=2), encoding="utf-8")


if __name__ == "__main__":
    main()
Evaluation ContractCategory, tested runtime, effort, and required artifacts+
Public identifierevaluation/blackout-cascade
CategoryGames Puzzles and Interactive Simulation
SubcategoryInteractive text games
Expert estimate4.0 hours
Model testedGPT-5.4
Agent testedTerminus-2
Required artifacts/app/solve.py, /app/output.json
Difficulty DesignWhy the benchmark discriminates between plausible and correct work+

Blackout Cascade extends the partizan game Blackout (Burke, Ferland, Huntemann, Teng, FUN 2024 / TCS 2025) with a nested adversarial control decision: a mover may retain direct control or irreversibly fracture a table, after which the mover nominates a switch pair while the opponent chooses the consumed switch and action. Correct play requires memoized game-tree reasoning with outcome-distance preferences applied separately to the mover and responder; always fracturing, never fracturing, and flattening nomination into ordinary play are distinct but plausible mistakes. A pending blackout, initiative, two-table switch fatigue, and XOR light state cross table boundaries, so local errors change later legal moves and scores. The supplied relay and verifier-only bank are deterministic synthetic positions: bounded bitmasks and persistent ids are procedurally generated, supplemented by hand-constructed carry, score-tie, and delayed-fracture probes. Combinatorial-game researchers and interactive rules-engine developers use optimal replay certificates like the requested artifact to validate a new solver or a compatibility migration against a reference semantics.

Reference ApproachThe intended solution strategy+

Parse positions.txt into each table's light mask, switch masks, and ids. Memoize each table state over (lights_mask, remaining_alloff_mask, remaining_oneon_mask, turn, blackout_run, fractured). In intact states compare ordinary switch/action choices with the free transition to fractured play, preferring intact on an exact tie. In fractured play enumerate legal one- or two-switch nominations; for each nomination let the opponent choose the consumed switch and action using that opponent's outcome-distance preference, then let the mover choose among the resolved nominations. Count only consumption or a pass in distance. Across tables, carry a pending blackout and next starter, enforce id-based fatigue, and XOR inherited lights. Replay the selected principal variation, award board-control points to each table winner, and serialize the scores, carry state, resource counts, and exact move line as the certificate.

Verification DesignHow the produced artifacts are independently checked+

The verifier checks both required artifacts, the complete nested replay-certificate schema, standard-library-only execution, and exact visible output. It validates the hidden expected JSON and deterministic compressed relay bank against SHA-256 sidecar entries, then reruns solve.py on every hidden relay and compares the complete certificate exactly. Mutation-selected relays reject never or immediate fracture, one-switch nominations, mover-friendly opponent responses, counting fracture as a move, outcome-only play, incorrect index/action tie-breaks, broken blackout or initiative carry, and reversed score or relay tie logic. Exact equality is appropriate because the rules, tie-breaks, inputs, and integer scoring are deterministic.

Task 09 · End-to-end software build

Reconcile Action Calendar

Agentic software engineering · Regulated Knowledge Work and Business Operations · Personal Assistant Productivity

Reconcile a versioned directive ledger into an optimal, policy-compliant executive action calendar.

Full brief · 5 source files · 732 lines · 4.0 expert hours

Task BriefThe complete agent-facing instruction+

Reconcile the versioned directive ledger in /app/directives.jsonl into the globally optimal executive action calendar required by the normative rules in /app/policy.md. Use the staff and delegation records in /app/staff.json, the deadline calendars in /app/business_calendars.json, the source events in /app/anchors.json, and the availability, commitments, travel times, and existing bookings in /app/calendar.json.

Write the complete audit trail, reconciled action states, optimal placements, deferral reasons, and objective totals to /app/reconciled_plan.json. The normative structure, field order, value constraints, serialization, and timestamp format for that file are defined in /app/output_schema.json.

Implementation & Verification Code5 authored source files: solution, public tests, and build environment+
Source Browsersolution/solve.py
5 files · 732 lines
#!/usr/bin/env python3
from __future__ import annotations

import copy
import json
from datetime import date, datetime, time, timedelta, timezone
from pathlib import Path


APP = Path("/app")
UTC = timezone.utc


def load_json(name):
    return json.loads((APP / name).read_text(encoding="utf-8"))


def parse_ts(value):
    return datetime.fromisoformat(value.replace("Z", "+00:00"))


def fmt_ts(value):
    return value.astimezone(UTC).strftime("%Y-%m-%dT%H:%M:%SZ")


staff = load_json("staff.json")
calendars = load_json("business_calendars.json")
anchors = load_json("anchors.json")
calendar = load_json("calendar.json")
rank = staff["classification_rank"]


def delegation_authorizes(issuer, role, scope, classification, instant):
    if staff["roles"][role] == issuer:
        return True
    for delegation in staff["delegations"]:
        if delegation["role"] != role or delegation["delegate"] != issuer:
            continue
        if not (parse_ts(delegation["start"]) <= instant < parse_ts(delegation["end"])):
            continue
        if scope not in delegation["scopes"]:
            continue
        if rank[delegation["max_classification"]] < rank[classification]:
            continue
        return True
    return False


def replay_ledger():
    records = [
        json.loads(line)
        for line in (APP / "directives.jsonl").read_text(encoding="utf-8").splitlines()
        if line.strip()
    ]
    records.sort(key=lambda item: item["sequence"])
    actions = {}
    accepted = []
    rejected = []

    for record in records:
        action_id = record["action_id"]
        kind = record["kind"]
        existing = actions.get(action_id)

        if kind == "open" and existing is not None:
            rejected.append({"event_id": record["event_id"], "reason": "duplicate_open"})
            continue
        if kind != "open" and existing is None:
            rejected.append({"event_id": record["event_id"], "reason": "missing_action"})
            continue
        if kind != "open" and record["base_version"] != existing["version"]:
            rejected.append({"event_id": record["event_id"], "reason": "stale_version"})
            continue

        instant = parse_ts(record["issued_at"])
        if kind == "open":
            fields = copy.deepcopy(record["fields"])
            authorized = delegation_authorizes(
                record["issuer"],
                record["controlling_role"],
                fields["scope"],
                fields["classification"],
                instant,
            )
        elif kind == "amend":
            fields = copy.deepcopy(existing["fields"])
            fields.update(record["patch"])
            old_ok = delegation_authorizes(
                record["issuer"],
                existing["controlling_role"],
                existing["fields"]["scope"],
                existing["fields"]["classification"],
                instant,
            )
            new_ok = delegation_authorizes(
                record["issuer"],
                existing["controlling_role"],
                fields["scope"],
                fields["classification"],
                instant,
            )
            authorized = old_ok and new_ok
        else:
            fields = existing["fields"]
            authorized = delegation_authorizes(
                record["issuer"],
                existing["controlling_role"],
                fields["scope"],
                fields["classification"],
                instant,
            )

        if not authorized:
            rejected.append({"event_id": record["event_id"], "reason": "unauthorized"})
            continue

        if kind == "open":
            actions[action_id] = {
                "action_id": action_id,
                "controlling_role": record["controlling_role"],
                "fields": fields,
                "version": 1,
                "active": True,
            }
        else:
            existing["version"] += 1
            if kind == "amend":
                existing["fields"] = fields
            elif kind == "cancel":
                existing["active"] = False
            elif kind == "reopen":
                existing["active"] = True
        accepted.append(record["event_id"])

    return actions, accepted, rejected


def is_open_day(day, calendar_id):
    return (
        day.weekday() < 5
        and day.isoformat() not in calendars[calendar_id]["closed_dates"]
    )


def next_open_day(day, calendar_id):
    candidate = day + timedelta(days=1)
    while not is_open_day(candidate, calendar_id):
        candidate += timedelta(days=1)
    return candidate


def deadline_for(fields):
    calendar_id = fields["deadline_calendar"]
    offset_text = calendars[calendar_id]["utc_offset"]
    sign = -1 if offset_text.startswith("-") else 1
    offset_hour, offset_minute = (int(part) for part in offset_text[1:].split(":"))
    zone = timezone(sign * timedelta(hours=offset_hour, minutes=offset_minute))
    anchor_local = parse_ts(anchors[fields["anchor_id"]]["end"]).astimezone(zone)
    if is_open_day(anchor_local.date(), calendar_id) and anchor_local.time() <= time(13, 0):
        business_zero = anchor_local.date()
    else:
        business_zero = next_open_day(anchor_local.date(), calendar_id)

    due_day = business_zero
    for _ in range(fields["business_days"]):
        due_day = next_open_day(due_day, calendar_id)
    hour, minute = (int(part) for part in fields["due_time"].split(":"))
    return datetime.combine(due_day, time(hour, minute), tzinfo=zone).astimezone(UTC)


def resolve_participants(fields, start, end):
    resolved = []
    for role in fields["required_roles"]:
        person = staff["roles"][role]
        for delegation in staff["delegations"]:
            if delegation["role"] != role:
                continue
            if not (
                parse_ts(delegation["start"]) <= start
                and end <= parse_ts(delegation["end"])
            ):
                continue
            if fields["scope"] not in delegation["scopes"]:
                continue
            if rank[delegation["max_classification"]] < rank[fields["classification"]]:
                continue
            person = delegation["delegate"]
            break
        if person not in resolved:
            resolved.append(person)
    return sorted(resolved)


def inside_availability(person, start, end):
    return any(
        parse_ts(window_start) <= start and end <= parse_ts(window_end)
        for window_start, window_end in calendar["availability"][person]
    )


def agenda_is_valid(appointments):
    ordered = sorted(appointments, key=lambda item: (item["start"], item["end"]))
    for left, right in zip(ordered, ordered[1:]):
        if left["end"] > right["start"]:
            return False
        travel = calendar["travel_minutes"][left["location"]][right["location"]]
        if left["end"] + timedelta(minutes=travel) > right["start"]:
            return False
    return True


fixed = {
    person: [
        {
            "start": parse_ts(item["start"]),
            "end": parse_ts(item["end"]),
            "location": item["location"],
        }
        for item in items
    ]
    for person, items in calendar["commitments"].items()
}


def candidate_is_standalone_valid(fields, start, end, participants):
    for person in participants:
        if rank[staff["people"][person]["clearance"]] < rank[fields["classification"]]:
            return False
        if not inside_availability(person, start, end):
            return False
        meeting = {"start": start, "end": end, "location": fields["location"]}
        if not agenda_is_valid(fixed[person] + [meeting]):
            return False
    return True


def disruption_cost(action_id, candidate):
    existing = calendar["existing_bookings"].get(action_id)
    if existing is None:
        return 1
    if (
        existing["start"] == fmt_ts(candidate["start"])
        and existing["end"] == fmt_ts(candidate["end"])
        and existing["location"] == candidate["location"]
        and sorted(existing["participants"]) == candidate["participants"]
    ):
        return 0
    return 2


def build_candidates(actions):
    planning_start = parse_ts(calendar["planning_start"])
    planning_end = parse_ts(calendar["planning_end"])
    step = timedelta(minutes=calendar["grid_minutes"])
    candidates = {}
    deadlines = {}

    for action_id, action in actions.items():
        if not action["active"]:
            continue
        fields = action["fields"]
        deadline = deadline_for(fields)
        deadlines[action_id] = deadline
        duration = timedelta(minutes=fields["duration_minutes"])
        end_limit = min(deadline, planning_end)
        options = []
        start = planning_start
        while start + duration <= end_limit:
            end = start + duration
            participants = resolve_participants(fields, start, end)
            if candidate_is_standalone_valid(fields, start, end, participants):
                candidate = {
                    "start": start,
                    "end": end,
                    "participants": participants,
                    "location": fields["location"],
                }
                candidate["cost"] = disruption_cost(action_id, candidate)
                options.append(candidate)
            start += step
        options.sort(
            key=lambda option: (
                option["cost"],
                option["start"],
                option["participants"],
            )
        )
        candidates[action_id] = options
    return candidates, deadlines


def topological_order(actions, candidates):
    active = {key for key, action in actions.items() if action["active"]}
    remaining = set(active)
    order = []
    completed = set()
    while remaining:
        ready = [
            action_id
            for action_id in remaining
            if set(actions[action_id]["fields"]["dependencies"]) <= completed
        ]
        if not ready:
            raise RuntimeError("dependency cycle")
        ready.sort(key=lambda item: (len(candidates[item]), -actions[item]["fields"]["priority"], item))
        chosen = ready[0]
        order.append(chosen)
        completed.add(chosen)
        remaining.remove(chosen)
    return order


def solve_schedule(actions, candidates):
    order = topological_order(actions, candidates)
    planning_start = parse_ts(calendar["planning_start"])
    maximum_remaining_priority = [0] * (len(order) + 1)
    maximum_remaining_count = [0] * (len(order) + 1)
    minimum_remaining_cost = [0] * (len(order) + 1)
    minimum_remaining_start = [0] * (len(order) + 1)
    for index in range(len(order) - 1, -1, -1):
        action_id = order[index]
        maximum_remaining_priority[index] = (
            maximum_remaining_priority[index + 1]
            + actions[action_id]["fields"]["priority"]
        )
        maximum_remaining_count[index] = maximum_remaining_count[index + 1] + 1
        if candidates[action_id]:
            minimum_remaining_cost[index] = (
                minimum_remaining_cost[index + 1]
                + min(option["cost"] for option in candidates[action_id])
            )
            minimum_remaining_start[index] = (
                minimum_remaining_start[index + 1]
                + min(
                    int((option["start"] - planning_start).total_seconds() // 60)
                    for option in candidates[action_id]
                )
            )
        else:
            existing_penalty = 3 if action_id in calendar["existing_bookings"] else 0
            minimum_remaining_cost[index] = minimum_remaining_cost[index + 1] + existing_penalty
            minimum_remaining_start[index] = minimum_remaining_start[index + 1]

    best_key = None
    best_assignment = None
    assignment = {}
    scheduled_by_person = {person: [] for person in staff["people"]}

    def compatible(candidate):
        appointment = {
            "start": candidate["start"],
            "end": candidate["end"],
            "location": candidate["location"],
        }
        for person in candidate["participants"]:
            if not agenda_is_valid(fixed[person] + scheduled_by_person[person] + [appointment]):
                return False
        return True

    def record_candidate(index, priority, count, cost, start_sum):
        nonlocal best_key, best_assignment
        placement = tuple(
            fmt_ts(assignment[action_id]["start"])
            if assignment.get(action_id) is not None
            else "~"
            for action_id in sorted(order)
        )
        key = (-priority, -count, cost, start_sum, placement)
        if best_key is None or key < best_key:
            best_key = key
            best_assignment = copy.deepcopy(assignment)

    def search(index, priority, count, cost, start_sum):
        if index == len(order):
            record_candidate(index, priority, count, cost, start_sum)
            return

        if best_key is not None:
            best_priority = -best_key[0]
            priority_ceiling = priority + maximum_remaining_priority[index]
            if priority_ceiling < best_priority:
                return
            if priority_ceiling == best_priority:
                best_count = -best_key[1]
                count_ceiling = count + maximum_remaining_count[index]
                if count_ceiling < best_count:
                    return
                if count_ceiling == best_count:
                    cost_floor = cost + minimum_remaining_cost[index]
                    if cost_floor > best_key[2]:
                        return
                    if cost_floor == best_key[2]:
                        start_floor = start_sum + minimum_remaining_start[index]
                        if start_floor > best_key[3]:
                            return

        action_id = order[index]
        fields = actions[action_id]["fields"]
        dependencies = fields["dependencies"]
        dependencies_scheduled = all(assignment.get(dep) is not None for dep in dependencies)

        if dependencies_scheduled:
            dependency_end = max(
                (assignment[dep]["end"] for dep in dependencies),
                default=parse_ts(calendar["planning_start"]),
            )
            for option in candidates[action_id]:
                if option["start"] < dependency_end:
                    continue
                if not compatible(option):
                    continue
                assignment[action_id] = option
                appointment = {
                    "start": option["start"],
                    "end": option["end"],
                    "location": option["location"],
                }
                for person in option["participants"]:
                    scheduled_by_person[person].append(appointment)
                search(
                    index + 1,
                    priority + fields["priority"],
                    count + 1,
                    cost + option["cost"],
                    start_sum
                    + int((option["start"] - planning_start).total_seconds() // 60),
                )
                for person in option["participants"]:
                    scheduled_by_person[person].pop()

        assignment[action_id] = None
        defer_cost = 3 if action_id in calendar["existing_bookings"] else 0
        search(index + 1, priority, count, cost + defer_cost, start_sum)
        del assignment[action_id]

    search(0, 0, 0, 0, 0)
    if best_assignment is None:
        raise RuntimeError("no schedule was evaluated")
    objective = {
        "scheduled_priority": -best_key[0],
        "scheduled_count": -best_key[1],
        "disruption_cost": best_key[2],
        "start_minute_sum": best_key[3],
    }
    return best_assignment, objective


def build_output():
    actions, accepted, rejected = replay_ledger()
    candidates, deadlines = build_candidates(actions)
    assignment, objective = solve_schedule(actions, candidates)
    output_actions = []

    for action_id in sorted(actions):
        action = actions[action_id]
        if not action["active"]:
            output_actions.append(
                {
                    "action_id": action_id,
                    "version": action["version"],
                    "status": "cancelled",
                    "deadline_utc": None,
                    "start_utc": None,
                    "end_utc": None,
                    "participants": [],
                    "reason": "cancelled",
                }
            )
            continue

        option = assignment[action_id]
        if option is None:
            blocked = any(assignment.get(dep) is None for dep in action["fields"]["dependencies"])
            output_actions.append(
                {
                    "action_id": action_id,
                    "version": action["version"],
                    "status": "deferred",
                    "deadline_utc": fmt_ts(deadlines[action_id]),
                    "start_utc": None,
                    "end_utc": None,
                    "participants": [],
                    "reason": "blocked" if blocked else "capacity",
                }
            )
        else:
            output_actions.append(
                {
                    "action_id": action_id,
                    "version": action["version"],
                    "status": "scheduled",
                    "deadline_utc": fmt_ts(deadlines[action_id]),
                    "start_utc": fmt_ts(option["start"]),
                    "end_utc": fmt_ts(option["end"]),
                    "participants": option["participants"],
                    "reason": None,
                }
            )

    return {
        "accepted_events": accepted,
        "rejected_events": rejected,
        "actions": output_actions,
        "objective": objective,
    }


result = build_output()
(APP / "reconciled_plan.json").write_text(
    json.dumps(result, indent=2) + "\n",
    encoding="utf-8",
)
Evaluation ContractCategory, tested runtime, effort, and required artifacts+
Public identifierevaluation/reconcile-action-calendar
CategoryRegulated Knowledge Work and Business Operations
SubcategoryPersonal Assistant Productivity
Expert estimate4.0 hours
Model testedOpus-4.8
Agent testedTerminus-2
Required artifacts/app/reconciled_plan.json
Difficulty DesignWhy the benchmark discriminates between plausible and correct work+

This task combines event-sourced record reconciliation, delegated authority, regional business-calendar arithmetic, dynamic role resolution, and constrained calendar optimization. The central trap is methodological consistency: a locally plausible decision about a stale amendment or delegate changes downstream action versions, deadlines, participants, dependencies, and the global optimum. The records and Aster Vale policy are synthetic but model realistic executive-office and regulated-operations workflows in which assistants must reconcile corrections, authorization windows, closures, fixed commitments, travel, and disruption costs before publishing an auditable action plan.

Reference ApproachThe intended solution strategy+

The reference solution replays records in policy order while validating versions and authorization against current and proposed action state. It derives each active deadline with the relevant offset and closure calendar, generates 30-minute candidate placements, resolves delegates for each full interval, and rejects candidates that violate clearance, availability, commitments, deadlines, travel, or dependency order. A branch-and-bound search then selects placements or deferrals using the policy's five-level lexicographic objective and serializes the audit trail and optimal action plan.

Verification DesignHow the produced artifacts are independently checked+

The verifier parses only /app/reconciled_plan.json and checks the normative schema, exact accepted and rejected event audit trail, resulting versions and lifecycle states, every derived business deadline, all optimal placements with their time-dependent participants and deferral reasons, and the four reported objective totals. The hidden expected plan is derived by the reference solver from the synthetic inputs; Nop cannot pass because the required artifact is absent, and non-optimal or locally inconsistent plans fail targeted assertions.

Task 10 · End-to-end software build

Archive Publisher

Agentic software engineering · Build Dependency and Release Management · Package publishing

Reconstruct a retired package publisher from surviving repository archives.

Full brief · 5 source files · 1,121 lines · 6 expert hours

Task BriefThe complete agent-facing instruction+

Reimplement the retired package publisher represented by the catalog-to-bundle records in /app/data/archives.json.

Write a self-contained program at /app/publisher.py with this interface:

python /app/publisher.py --catalog CATALOG --output OUTPUT

CATALOG and OUTPUT will be absolute paths. The catalog format, field semantics, exact bundle schema, and serialization contract are in /app/data/ARCHIVES.md; treat that file and the archived records as normative. Your program will be evaluated on other valid catalogs that combine the same archived behaviors.

/app/data/check_archives.py runs the replacement against every surviving record. Do not hardcode archive IDs, project names, versions, filenames, digests, or results.

Run the publisher on /app/data/sample.json and write the resulting bundle to /app/output. It must correspond to /app/data/sample-expected.json.

A successful run must replace OUTPUT with a complete bundle containing no stale files. Malformed or unpublishable catalogs must exit nonzero without changing an existing output directory. Do not modify anything under /app/data/.

Implementation & Verification Code5 authored source files: solution, public tests, and build environment+
Source Browsersolution/archive_publisher.py
5 files · 1,121 lines
#!/usr/bin/env python3
from __future__ import annotations

import argparse
import itertools
import json
import os
import re
import shutil
import sys
import tempfile
from pathlib import Path
from typing import Any


class PublishError(ValueError):
    pass


def canonical(value: object) -> bytes:
    return json.dumps(
        value,
        ensure_ascii=False,
        sort_keys=True,
        separators=(",", ":"),
    ).encode("utf-8")


def require_int(value: object, label: str, minimum: int) -> int:
    if isinstance(value, bool) or not isinstance(value, int) or value < minimum:
        raise PublishError(f"{label} must be an integer >= {minimum}")
    return value


def require_name(value: object, label: str) -> str:
    if not isinstance(value, str) or not re.fullmatch(r"[a-z][a-z0-9-]*", value):
        raise PublishError(f"{label} is invalid")
    return value


def parse_catalog(path: Path) -> dict[str, Any]:
    try:
        raw = json.loads(path.read_text(encoding="utf-8"))
    except (OSError, UnicodeError, json.JSONDecodeError) as exc:
        raise PublishError(f"cannot read catalog: {exc}") from exc
    if not isinstance(raw, dict):
        raise PublishError("catalog must be an object")

    lane_count = require_int(raw.get("lane_count"), "lane_count", 1)
    bytes_per_tick = require_int(
        raw.get("bytes_per_tick"), "bytes_per_tick", 1
    )
    projects = raw.get("projects")
    if not isinstance(projects, list) or not projects:
        raise PublishError("projects must be a nonempty array")

    names: list[str] = []
    for index, project in enumerate(projects):
        if not isinstance(project, dict):
            raise PublishError(f"projects[{index}] must be an object")
        name = require_name(project.get("name"), f"projects[{index}].name")
        if name in names:
            raise PublishError(f"duplicate project {name}")
        names.append(name)

    known = set(names)
    chunk_sizes: dict[str, int] = {}
    vector_count = 1
    for project_index, project in enumerate(projects):
        name = names[project_index]
        candidates = project.get("candidates")
        if not isinstance(candidates, list) or not candidates:
            raise PublishError(f"{name}.candidates must be nonempty")
        vector_count *= len(candidates)
        if vector_count > 200_000:
            raise PublishError("candidate vector space exceeds 200000")

        versions: set[str] = set()
        for candidate_index, candidate in enumerate(candidates):
            label = f"{name}.candidates[{candidate_index}]"
            if not isinstance(candidate, dict):
                raise PublishError(f"{label} must be an object")
            version = candidate.get("version")
            if not isinstance(version, str) or not version or version in versions:
                raise PublishError(f"{label}.version is invalid or duplicated")
            versions.add(version)
            require_int(candidate.get("signed_until"), f"{label}.signed_until", 0)

            filename = candidate.get("filename")
            if (
                not isinstance(filename, str)
                or not filename
                or "/" in filename
                or "\\" in filename
            ):
                raise PublishError(f"{label}.filename is invalid")
            digest = candidate.get("sha256")
            if not isinstance(digest, str) or not re.fullmatch(
                r"[0-9a-f]{64}", digest
            ):
                raise PublishError(f"{label}.sha256 is invalid")

            for field in ("requires", "atomic_with"):
                targets = candidate.get(field)
                if (
                    not isinstance(targets, list)
                    or any(not isinstance(target, str) for target in targets)
                    or len(targets) != len(set(targets))
                ):
                    raise PublishError(f"{label}.{field} is invalid")
                for target in targets:
                    if target not in known or target == name:
                        raise PublishError(f"{label}.{field} has invalid target")

            chunks = candidate.get("chunks")
            if not isinstance(chunks, list) or not chunks:
                raise PublishError(f"{label}.chunks must be nonempty")
            seen_chunks: set[str] = set()
            for chunk_index, chunk in enumerate(chunks):
                chunk_label = f"{label}.chunks[{chunk_index}]"
                if not isinstance(chunk, dict):
                    raise PublishError(f"{chunk_label} must be an object")
                chunk_id = chunk.get("id")
                if (
                    not isinstance(chunk_id, str)
                    or not re.fullmatch(r"[a-z0-9][a-z0-9-]*", chunk_id)
                    or chunk_id in seen_chunks
                ):
                    raise PublishError(f"{chunk_label}.id is invalid")
                seen_chunks.add(chunk_id)
                size = require_int(chunk.get("bytes"), f"{chunk_label}.bytes", 1)
                prior = chunk_sizes.setdefault(chunk_id, size)
                if prior != size:
                    raise PublishError(f"chunk {chunk_id} has inconsistent size")

    return {
        "lane_count": lane_count,
        "bytes_per_tick": bytes_per_tick,
        "projects": projects,
        "names": names,
        "chunk_sizes": chunk_sizes,
    }


class UnionFind:
    def __init__(self, size: int) -> None:
        self.parent = list(range(size))

    def find(self, value: int) -> int:
        while self.parent[value] != value:
            self.parent[value] = self.parent[self.parent[value]]
            value = self.parent[value]
        return value

    def union(self, left: int, right: int) -> None:
        left_root = self.find(left)
        right_root = self.find(right)
        if left_root != right_root:
            self.parent[right_root] = left_root


def schedule_vector(
    data: dict[str, Any],
    indices: tuple[int, ...],
) -> tuple[dict[str, Any], list[dict[str, Any]]] | None:
    projects = data["projects"]
    names = data["names"]
    position = {name: index for index, name in enumerate(names)}
    selected = [
        project["candidates"][candidate_index]
        for project, candidate_index in zip(projects, indices)
    ]

    groups = UnionFind(len(projects))
    for index, candidate in enumerate(selected):
        for target in candidate["atomic_with"]:
            groups.union(index, position[target])

    components_by_root: dict[int, list[int]] = {}
    for index in range(len(projects)):
        components_by_root.setdefault(groups.find(index), []).append(index)
    components = sorted(
        components_by_root.values(),
        key=lambda members: members[0],
    )
    component_of: dict[int, int] = {}
    for component_index, members in enumerate(components):
        for member in members:
            component_of[member] = component_index

    predecessors = [set() for _ in components]
    for dependent, candidate in enumerate(selected):
        for requirement in candidate["requires"]:
            dependency = position[requirement]
            left = component_of[dependency]
            right = component_of[dependent]
            if left != right:
                predecessors[right].add(left)

    component_ids = [
        "+".join(names[index] for index in members)
        for members in components
    ]
    lane_clocks = [0] * data["lane_count"]
    reserved: dict[str, int] = {}
    complete_by_component: dict[int, int] = {}
    scheduled: set[int] = set()
    scheduled_order: list[int] = []
    records: list[dict[str, Any]] = []

    while len(scheduled) < len(components):
        ready = [
            index
            for index in range(len(components))
            if index not in scheduled and predecessors[index] <= scheduled
        ]
        if not ready:
            return None

        lane = min(range(len(lane_clocks)), key=lambda item: (lane_clocks[item], item))
        projections: list[tuple[int, int, dict[str, Any]]] = []
        for component_index in ready:
            members = components[component_index]
            dependency_ready = max(
                (complete_by_component[item] for item in predecessors[component_index]),
                default=0,
            )
            chunks: dict[str, int] = {}
            for member in members:
                for chunk in selected[member]["chunks"]:
                    chunks[chunk["id"]] = chunk["bytes"]
            charged = sorted(chunk for chunk in chunks if chunk not in reserved)
            chunk_ready = max(
                (reserved[chunk] for chunk in chunks if chunk in reserved),
                default=0,
            )
            start = max(lane_clocks[lane], dependency_ready, chunk_ready)
            uploaded_bytes = sum(chunks[chunk] for chunk in charged)
            duration = (
                uploaded_bytes + data["bytes_per_tick"] - 1
            ) // data["bytes_per_tick"]
            complete = start + duration
            signed_until = min(selected[member]["signed_until"] for member in members)
            headroom = signed_until - complete
            projection = {
                "charged_chunks": charged,
                "complete_at": complete,
                "headroom": headroom,
                "id": component_ids[component_index],
                "lane": lane,
                "projects": [names[member] for member in members],
                "start_at": start,
                "uploaded_bytes": uploaded_bytes,
                "wait_for": sorted(
                    component_ids[item] for item in predecessors[component_index]
                ),
            }
            projections.append((headroom, members[0], projection))

        _, _, record = min(projections, key=lambda item: (item[0], item[1]))
        component_index = component_ids.index(record["id"])
        scheduled.add(component_index)
        scheduled_order.append(component_index)
        complete_by_component[component_index] = record["complete_at"]
        lane_clocks[lane] = record["complete_at"]
        for chunk in record["charged_chunks"]:
            reserved[chunk] = record["complete_at"]
        records.append(record)

    warranted_through = dict(complete_by_component)
    for component_index in reversed(scheduled_order):
        for predecessor in predecessors[component_index]:
            warranted_through[predecessor] = max(
                warranted_through[predecessor],
                warranted_through[component_index],
            )
    for component_index, record in zip(scheduled_order, records):
        signed_until = min(
            selected[member]["signed_until"]
            for member in components[component_index]
        )
        record["headroom"] = (
            signed_until - warranted_through[component_index]
        )

    if any(record["headroom"] < 0 for record in records):
        return None

    manifest = {
        "cohorts": records,
        "published_at": max(record["complete_at"] for record in records),
        "selected": {
            name: candidate["version"]
            for name, candidate in zip(names, selected)
        },
    }
    return manifest, selected


def build_bundle(data: dict[str, Any]) -> dict[str, Any]:
    ranges = [
        range(len(project["candidates"])) for project in data["projects"]
    ]
    result: tuple[dict[str, Any], list[dict[str, Any]]] | None = None
    for indices in itertools.product(*ranges):
        result = schedule_vector(data, indices)
        if result is not None:
            break
    if result is None:
        raise PublishError("catalog has no publishable candidate vector")

    manifest, selected = result
    pages: dict[str, Any] = {}
    for name, candidate in zip(data["names"], selected):
        chunks = sorted(candidate["chunks"], key=lambda chunk: chunk["id"])
        pages[name] = {
            "files": [{
                "chunks": [chunk["id"] for chunk in chunks],
                "filename": candidate["filename"],
                "hashes": {"sha256": candidate["sha256"]},
                "size": sum(chunk["bytes"] for chunk in chunks),
            }],
            "meta": {"api-version": "1.0"},
            "name": name,
            "requires": sorted(candidate["requires"]),
            "version": candidate["version"],
        }
    return {"manifest": manifest, "pages": pages}


def replace_bundle(path: Path, bundle: dict[str, Any]) -> None:
    parent = path.parent
    if not parent.is_dir():
        raise PublishError("output parent does not exist")
    temporary = Path(tempfile.mkdtemp(prefix=f".{path.name}.", dir=parent))
    try:
        (temporary / "simple").mkdir()
        (temporary / "manifest.json").write_bytes(canonical(bundle["manifest"]))
        for name, page in bundle["pages"].items():
            (temporary / "simple" / f"{name}.json").write_bytes(canonical(page))
        if path.exists():
            if path.is_symlink() or not path.is_dir():
                raise PublishError("existing output is not a directory")
            shutil.rmtree(path)
        os.replace(temporary, path)
    except BaseException:
        shutil.rmtree(temporary, ignore_errors=True)
        raise


def main() -> int:
    parser = argparse.ArgumentParser()
    parser.add_argument("--catalog", required=True)
    parser.add_argument("--output", required=True)
    args = parser.parse_args()
    try:
        data = parse_catalog(Path(args.catalog))
        bundle = build_bundle(data)
        replace_bundle(Path(args.output), bundle)
    except (OSError, PublishError) as exc:
        print(f"archive publisher failed: {exc}", file=sys.stderr)
        return 1
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
Evaluation ContractCategory, tested runtime, effort, and required artifacts+
Public identifierevaluation/archive-publisher
CategoryBuild Dependency and Release Management
SubcategoryPackage publishing
Expert estimate6 hours
Model testedOpus-4.8
Agent testedTerminus-2
Required artifacts/app/publisher.py, /app/output
Difficulty DesignWhy the benchmark discriminates between plausible and correct work+

A release-infrastructure engineer must recover a deterministic internal publisher from compact catalog-to-repository archives. The archive corpus is synthetic but constructed to concentrate realistic release-recovery constraints into reviewable examples. Candidate choice reshapes atomic cohorts and dependency edges; each publication changes lane clocks and reusable-chunk availability; and the completed schedule then propagates a dependency warranty backward, so a late dependent can invalidate an apparently healthy earlier artifact and force the entire vector to be rebuilt. The task is compact, but its scheduling, graph, and fallback calculations are coupled and must remain correct under recomputation.

Reference ApproachThe intended solution strategy+

Validate the catalog, then consider candidate vectors in project-priority order. For each vector, form undirected atomic cohorts, derive inter-cohort dependencies, and schedule ready cohorts on the earliest lane by least projected upload slack. Charge each content chunk once and make later users wait for its reservation. After the complete schedule is known, propagate the latest dependent completion backward through the cohort DAG and recompute signing headroom. Reject the vector if any warranty is negative; otherwise emit the canonical manifest and project pages, replacing the output only after success.

Verification DesignHow the produced artifacts are independently checked+

Pytest integrity-checks the visible archive corpus, runs the submitted CLI unprivileged in isolated directories, and checks exact canonical bundles, output preservation, and verifier isolation. Held-out catalogs combine candidate fallback, atomic topology changes, dependency chains, parallel lanes, reusable chunks, repeated frontier projection, and a late dependent whose warranty reaches backward and invalidates a locally feasible preferred vector.

Task 11 · End-to-end software build

Cohort Harmonizer

Agentic software engineering · Data Science and Reporting · Dataset preparation for analysis

Prepare a protocol-aligned cohort by resolving bitemporal deliveries, inferred decoders, ambiguous identities, atomic group constraints, and a global resource allocation objective.

Full brief · 5 source files · 923 lines · 5 expert hours

Task BriefThe complete agent-facing instruction+

The protocol-drift export under /app/data must be harmonized into a trustworthy analysis cohort. Build a reusable Python program at /app/harmonize.py that implements the normative input, transformation, optimization, and output contract in /app/data/SPEC.md.

Your program must accept an input directory and an output directory:

python3 /app/harmonize.py INPUT_DIR OUTPUT_DIR

It will be run on other directories that follow the same contract, so derive every result from the supplied files rather than from values specific to this export.

Run it once for the shipped export with:

python3 /app/harmonize.py /app/data /app/output

This must create /app/output/cohort.csv, /app/output/excluded.csv, /app/output/deferred_groups.csv, /app/output/lineage.jsonl, and /app/output/summary.json, with the schemas, ordering, point-in-time rules, exact-decimal semantics, global objective, and reconciliation requirements defined in /app/data/SPEC.md. Do not modify any file under /app/data.

Implementation & Verification Code5 authored source files: solution, public tests, and build environment+
Source Browsersolution/solve.py
5 files · 923 lines
#!/usr/bin/env python3
import csv
import itertools
import json
import sys
from collections import Counter, defaultdict
from decimal import Decimal
from fractions import Fraction
from pathlib import Path


EXCLUSION_REASONS = (
    "FUTURE",
    "LATE_INGEST",
    "SUPERSEDED",
    "DECODER_UNRESOLVED",
    "MISSING_TOKEN",
    "NO_LINK",
    "NO_SLOT",
    "NOT_SELECTED",
)

COHORT_FIELDS = [
    "group_id",
    "slot_id",
    "subject_id",
    "visit",
    "analyte",
    "measurement_id",
    "delivery_id",
    "batch_id",
    "decoder_id",
    "link_id",
    "canonical_value",
    "unit",
    "payload_id",
    "candidate_score",
]
EXCLUDED_FIELDS = ["delivery_id", "measurement_id", "reason"]
DEFERRED_FIELDS = ["group_id", "reason"]
AMENDABLE_FIELDS = {
    "batch_id",
    "local_subject",
    "visit",
    "analyte",
    "raw_value",
    "quality",
    "payload_id",
}


def read_csv(path):
    with path.open(newline="", encoding="utf-8") as handle:
        return list(csv.DictReader(handle))


def decimal_text(value):
    if value == 0:
        return "0"
    return format(value.normalize(), "f")


def write_csv(path, fieldnames, rows):
    with path.open("w", newline="", encoding="utf-8") as handle:
        writer = csv.DictWriter(handle, fieldnames=fieldnames)
        writer.writeheader()
        writer.writerows(rows)


def infer_decoders(profiles, controls, observation_cutoff, knowledge_cutoff):
    profiles_by_batch = defaultdict(list)
    for profile in profiles:
        profiles_by_batch[profile["batch_id"]].append(profile)

    eligible_by_batch = defaultdict(list)
    for control in controls:
        if (
            control["observed_at"] <= observation_cutoff
            and control["recorded_at"] <= knowledge_cutoff
        ):
            eligible_by_batch[control["batch_id"]].append(control)

    resolutions = {}
    control_ids = {}
    for batch_id in sorted(profiles_by_batch):
        eligible = eligible_by_batch[batch_id]
        control_ids[batch_id] = sorted(row["control_id"] for row in eligible)
        fitting = []
        for profile in profiles_by_batch[batch_id]:
            if not eligible:
                continue
            fits = True
            for control in eligible:
                if (
                    control["raw_value"] == profile["missing_token"]
                    or profile["output_unit"] != control["expected_unit"]
                    or Decimal(control["raw_value"]) * Decimal(profile["scale"])
                    + Decimal(profile["offset"])
                    != Decimal(control["expected_value"])
                ):
                    fits = False
                    break
            if fits:
                fitting.append(profile)
        resolutions[batch_id] = fitting[0] if len(fitting) == 1 else None
    return resolutions, control_ids


def select_deliveries(measurements, observation_cutoff, knowledge_cutoff):
    dispositions = {}
    eligible_by_measurement = defaultdict(list)
    cutoff_eligible = 0
    for row in measurements:
        if row["observed_at"] > observation_cutoff:
            dispositions[row["delivery_id"]] = "FUTURE"
        elif row["recorded_at"] > knowledge_cutoff:
            dispositions[row["delivery_id"]] = "LATE_INGEST"
        else:
            cutoff_eligible += 1
            eligible_by_measurement[row["measurement_id"]].append(row)

    winners = []
    for rows in eligible_by_measurement.values():
        winner = max(
            rows,
            key=lambda row: (
                row["recorded_at"],
                int(row["source_seq"]),
                row["delivery_id"],
            ),
        )
        winners.append(winner)
        for row in rows:
            if row is not winner:
                dispositions[row["delivery_id"]] = "SUPERSEDED"
    return winners, dispositions, cutoff_eligible


def apply_amendments(delivery, amendments, knowledge_cutoff):
    eligible_by_field = defaultdict(list)
    for amendment in amendments:
        if (
            amendment["measurement_id"] == delivery["measurement_id"]
            and amendment["field"] in AMENDABLE_FIELDS
            and amendment["effective_at"] <= delivery["observed_at"]
            and amendment["recorded_at"] <= knowledge_cutoff
        ):
            eligible_by_field[amendment["field"]].append(amendment)

    amended = dict(delivery)
    amendment_ids = []
    for field, candidates in eligible_by_field.items():
        winner = max(
            candidates,
            key=lambda row: (
                row["effective_at"],
                row["recorded_at"],
                int(row["revision"]),
                row["amendment_id"],
            ),
        )
        amended[field] = winner["new_value"]
        amendment_ids.append(winner["amendment_id"])
    return amended, sorted(amendment_ids)


def matching_links(delivery, links, knowledge_cutoff):
    return [
        link
        for link in links
        if link["batch_id"] == delivery["batch_id"]
        and link["local_subject"] == delivery["local_subject"]
        and link["valid_from"] <= delivery["observed_at"] < link["valid_to"]
        and link["effective_at"] <= delivery["observed_at"]
        and link["recorded_at"] <= knowledge_cutoff
    ]


def constraints_hold(constraints, values):
    for constraint in constraints:
        kind = constraint["type"]
        if kind == "sum_equals":
            if values[constraint["target"]] != sum(
                (values[slot] for slot in constraint["terms"]), Decimal(0)
            ):
                return False
        elif kind == "strictly_increasing":
            sequence = [values[slot] for slot in constraint["slots"]]
            if any(left >= right for left, right in zip(sequence, sequence[1:])):
                return False
        elif kind == "ratio_range":
            denominator = values[constraint["denominator"]]
            if denominator == 0:
                return False
            ratio = Fraction(values[constraint["numerator"]]) / Fraction(denominator)
            if not (
                Fraction(Decimal(constraint["min"]))
                <= ratio
                <= Fraction(Decimal(constraint["max"]))
            ):
                return False
        else:
            raise ValueError(f"unsupported constraint type: {kind}")
    return True


def payload_capacity(payload_id, protocol):
    return int(
        protocol["payload_capacity_overrides"].get(
            payload_id, protocol["default_payload_capacity"]
        )
    )


def enumerate_group_options(groups, candidates, protocol):
    options_by_group = {}
    for group in groups:
        candidates_by_slot = []
        for slot in group["slots"]:
            matches = [
                candidate
                for candidate in candidates
                if candidate["subject_id"] == slot["subject_id"]
                and candidate["visit"] == slot["visit"]
                and candidate["analyte"] == slot["analyte"]
                and candidate["unit"] == slot["unit"]
            ]
            candidates_by_slot.append(matches)

        options = []
        if all(candidates_by_slot):
            for combination in itertools.product(*candidates_by_slot):
                measurement_ids = [row["measurement_id"] for row in combination]
                if len(set(measurement_ids)) != len(measurement_ids):
                    continue
                payloads = Counter(row["payload_id"] for row in combination)
                if any(
                    count > payload_capacity(payload_id, protocol)
                    for payload_id, count in payloads.items()
                ):
                    continue
                values = {
                    slot["slot_id"]: candidate["value"]
                    for slot, candidate in zip(group["slots"], combination)
                }
                if not constraints_hold(group["constraints"], values):
                    continue
                rows = [
                    {"group": group, "slot": slot, "candidate": candidate}
                    for slot, candidate in zip(group["slots"], combination)
                ]
                options.append(
                    {
                        "rows": rows,
                        "measurements": set(measurement_ids),
                        "payloads": payloads,
                        "score": sum(row["candidate_score"] for row in combination),
                        "amendments": sum(
                            row["amendment_count"] for row in combination
                        ),
                    }
                )
        options.sort(
            key=lambda option: "|".join(
                f'{row["slot"]["slot_id"]}={row["candidate"]["candidate_id"]}'
                for row in sorted(
                    option["rows"], key=lambda row: row["slot"]["slot_id"]
                )
            )
        )
        options_by_group[group["group_id"]] = options
    return options_by_group


def optimize(groups, options_by_group, protocol):
    ordered_groups = sorted(groups, key=lambda group: group["group_id"])
    best = None

    def visit(
        index,
        used_measurements,
        used_payloads,
        selected_rows,
        priority_sum,
        group_count,
        score_sum,
        amendment_count,
    ):
        nonlocal best
        if index == len(ordered_groups):
            ordered_rows = sorted(
                selected_rows,
                key=lambda row: (
                    row["group"]["group_id"],
                    row["slot"]["slot_id"],
                ),
            )
            signature = "|".join(
                f'{row["slot"]["slot_id"]}={row["candidate"]["candidate_id"]}'
                for row in ordered_rows
            )
            numeric = (
                priority_sum,
                group_count,
                score_sum,
                -amendment_count,
            )
            if (
                best is None
                or numeric > best["numeric"]
                or (numeric == best["numeric"] and signature < best["signature"])
            ):
                best = {
                    "numeric": numeric,
                    "signature": signature,
                    "rows": list(ordered_rows),
                    "priority_sum": priority_sum,
                    "group_count": group_count,
                    "score_sum": score_sum,
                    "amendment_count": amendment_count,
                }
            return

        group = ordered_groups[index]
        visit(
            index + 1,
            used_measurements,
            used_payloads,
            selected_rows,
            priority_sum,
            group_count,
            score_sum,
            amendment_count,
        )
        for option in options_by_group[group["group_id"]]:
            if option["measurements"] & used_measurements:
                continue
            if any(
                used_payloads[payload_id] + count
                > payload_capacity(payload_id, protocol)
                for payload_id, count in option["payloads"].items()
            ):
                continue
            next_payloads = used_payloads + option["payloads"]
            visit(
                index + 1,
                used_measurements | option["measurements"],
                next_payloads,
                selected_rows + option["rows"],
                priority_sum + int(group["priority"]),
                group_count + 1,
                score_sum + option["score"],
                amendment_count + option["amendments"],
            )

    visit(0, set(), Counter(), [], 0, 0, 0, 0)
    return best


def harmonize(input_dir, output_dir):
    input_dir = Path(input_dir)
    output_dir = Path(output_dir)
    output_dir.mkdir(parents=True, exist_ok=True)

    cutoff = json.loads((input_dir / "cutoff.json").read_text(encoding="utf-8"))
    profiles = read_csv(input_dir / "profiles.csv")
    controls = read_csv(input_dir / "controls.csv")
    measurements = read_csv(input_dir / "measurements.csv")
    amendments = read_csv(input_dir / "amendments.csv")
    links = read_csv(input_dir / "links.csv")
    protocol = json.loads((input_dir / "protocol.json").read_text(encoding="utf-8"))

    observation_cutoff = cutoff["observation_cutoff"]
    knowledge_cutoff = cutoff["knowledge_cutoff"]
    resolutions, control_ids = infer_decoders(
        profiles, controls, observation_cutoff, knowledge_cutoff
    )
    winners, dispositions, cutoff_eligible = select_deliveries(
        measurements, observation_cutoff, knowledge_cutoff
    )

    all_slots = [
        slot for group in protocol["groups"] for slot in group["slots"]
    ]
    candidates = []
    matched_delivery_ids = set()
    prepared_by_delivery = {}

    for original in winners:
        amended, amendment_ids = apply_amendments(
            original, amendments, knowledge_cutoff
        )
        delivery_id = original["delivery_id"]
        prepared_by_delivery[delivery_id] = {
            "amended": amended,
            "amendment_ids": amendment_ids,
        }
        profile = resolutions.get(amended["batch_id"])
        if profile is None:
            dispositions[delivery_id] = "DECODER_UNRESOLVED"
            continue
        if amended["raw_value"] == profile["missing_token"]:
            dispositions[delivery_id] = "MISSING_TOKEN"
            continue

        value = Decimal(amended["raw_value"]) * Decimal(
            profile["scale"]
        ) + Decimal(profile["offset"])
        eligible_links = matching_links(amended, links, knowledge_cutoff)
        if not eligible_links:
            dispositions[delivery_id] = "NO_LINK"
            continue

        delivery_candidates = []
        for link in eligible_links:
            candidate = {
                "candidate_id": f'{delivery_id}@{link["link_id"]}',
                "delivery_id": delivery_id,
                "measurement_id": original["measurement_id"],
                "batch_id": amended["batch_id"],
                "decoder_id": profile["decoder_id"],
                "link_id": link["link_id"],
                "subject_id": link["subject_id"],
                "visit": amended["visit"],
                "analyte": amended["analyte"],
                "value": value,
                "unit": profile["output_unit"],
                "payload_id": amended["payload_id"],
                "candidate_score": int(protocol["quality_scores"][amended["quality"]])
                + int(link["score"]),
                "amendment_ids": amendment_ids,
                "amendment_count": len(amendment_ids),
            }
            delivery_candidates.append(candidate)

        if not any(
            candidate["subject_id"] == slot["subject_id"]
            and candidate["visit"] == slot["visit"]
            and candidate["analyte"] == slot["analyte"]
            and candidate["unit"] == slot["unit"]
            for candidate in delivery_candidates
            for slot in all_slots
        ):
            dispositions[delivery_id] = "NO_SLOT"
            continue
        matched_delivery_ids.add(delivery_id)
        candidates.extend(delivery_candidates)

    options_by_group = enumerate_group_options(
        protocol["groups"], candidates, protocol
    )
    optimum = optimize(protocol["groups"], options_by_group, protocol)
    selected_delivery_ids = {
        row["candidate"]["delivery_id"] for row in optimum["rows"]
    }
    for delivery_id in matched_delivery_ids - selected_delivery_ids:
        dispositions[delivery_id] = "NOT_SELECTED"

    cohort_rows = []
    lineage_rows = []
    completed_groups = set()
    for selected in optimum["rows"]:
        group = selected["group"]
        slot = selected["slot"]
        candidate = selected["candidate"]
        completed_groups.add(group["group_id"])
        cohort_rows.append(
            {
                "group_id": group["group_id"],
                "slot_id": slot["slot_id"],
                "subject_id": candidate["subject_id"],
                "visit": candidate["visit"],
                "analyte": candidate["analyte"],
                "measurement_id": candidate["measurement_id"],
                "delivery_id": candidate["delivery_id"],
                "batch_id": candidate["batch_id"],
                "decoder_id": candidate["decoder_id"],
                "link_id": candidate["link_id"],
                "canonical_value": decimal_text(candidate["value"]),
                "unit": candidate["unit"],
                "payload_id": candidate["payload_id"],
                "candidate_score": str(candidate["candidate_score"]),
            }
        )
        lineage_rows.append(
            {
                "group_id": group["group_id"],
                "slot_id": slot["slot_id"],
                "candidate_id": candidate["candidate_id"],
                "delivery_id": candidate["delivery_id"],
                "measurement_id": candidate["measurement_id"],
                "amendment_ids": candidate["amendment_ids"],
                "decoder_id": candidate["decoder_id"],
                "control_ids": control_ids[candidate["batch_id"]],
                "link_id": candidate["link_id"],
                "payload_id": candidate["payload_id"],
            }
        )

    cohort_rows.sort(key=lambda row: (row["group_id"], row["slot_id"]))
    lineage_rows.sort(key=lambda row: (row["group_id"], row["slot_id"]))
    excluded_rows = [
        {
            "delivery_id": row["delivery_id"],
            "measurement_id": row["measurement_id"],
            "reason": dispositions[row["delivery_id"]],
        }
        for row in measurements
        if row["delivery_id"] not in selected_delivery_ids
    ]
    excluded_rows.sort(key=lambda row: (row["measurement_id"], row["delivery_id"]))

    deferred_rows = []
    for group in protocol["groups"]:
        group_id = group["group_id"]
        if group_id not in completed_groups:
            deferred_rows.append(
                {
                    "group_id": group_id,
                    "reason": (
                        "NO_FEASIBLE_ASSIGNMENT"
                        if not options_by_group[group_id]
                        else "RESOURCE_CONFLICT"
                    ),
                }
            )
    deferred_rows.sort(key=lambda row: row["group_id"])

    write_csv(output_dir / "cohort.csv", COHORT_FIELDS, cohort_rows)
    write_csv(output_dir / "excluded.csv", EXCLUDED_FIELDS, excluded_rows)
    write_csv(output_dir / "deferred_groups.csv", DEFERRED_FIELDS, deferred_rows)
    with (output_dir / "lineage.jsonl").open("w", encoding="utf-8") as handle:
        for row in lineage_rows:
            handle.write(
                json.dumps(row, sort_keys=True, separators=(",", ":")) + "\n"
            )

    reason_counts = Counter(row["reason"] for row in excluded_rows)
    decoder_by_batch = {}
    for batch_id in sorted(resolutions):
        profile = resolutions[batch_id]
        decoder_by_batch[batch_id] = {
            "decoder_id": profile["decoder_id"] if profile else None,
            "control_ids": control_ids[batch_id],
        }
    summary = {
        "input_deliveries": len(measurements),
        "cutoff_eligible_deliveries": cutoff_eligible,
        "winning_deliveries": len(winners),
        "selected_deliveries": len(cohort_rows),
        "excluded_by_reason": {
            reason: reason_counts.get(reason, 0) for reason in EXCLUSION_REASONS
        },
        "decoder_by_batch": decoder_by_batch,
        "completed_groups": sorted(completed_groups),
        "deferred_groups": sorted(
            group["group_id"]
            for group in protocol["groups"]
            if group["group_id"] not in completed_groups
        ),
        "objective": {
            "priority_sum": optimum["priority_sum"],
            "group_count": optimum["group_count"],
            "candidate_score_sum": optimum["score_sum"],
            "amendment_field_count": optimum["amendment_count"],
            "selection_signature": optimum["signature"],
        },
    }
    (output_dir / "summary.json").write_text(
        json.dumps(summary, indent=2) + "\n", encoding="utf-8"
    )


def main():
    if len(sys.argv) != 3:
        raise SystemExit("usage: harmonize.py INPUT_DIR OUTPUT_DIR")
    harmonize(sys.argv[1], sys.argv[2])


if __name__ == "__main__":
    main()
Evaluation ContractCategory, tested runtime, effort, and required artifacts+
Public identifierevaluation/cohort-harmonizer
CategoryData Science and Reporting
SubcategoryDataset preparation for analysis
Expert estimate5 hours
Model testedOpus-4.8
Agent testedTerminus-2
Required artifacts/app/harmonize.py, /app/output/cohort.csv, /app/output/excluded.csv, /app/output/deferred_groups.csv, /app/output/lineage.jsonl, /app/output/summary.json
Difficulty DesignWhy the benchmark discriminates between plausible and correct work+

The task is a category-native dataset-preparation problem whose individually familiar operations interact nonlocally. A solution must infer batch decoders only from cutoff-valid controls, resolve bitemporal redeliveries and independent field amendments, preserve ambiguous point-in-time identity candidates, enumerate atomic protocol assignments under exact decimal constraints, and solve a lexicographic set-packing objective with shared measurement and payload capacities. Greedy decoder, link, candidate, or group choices all yield plausible but incorrect cohorts. The visible export contains misleading high-score alternatives, while the held-out export changes the conflict topology and independently forces the group-count, candidate-score, amendment-count, and final-signature levels. The hand-crafted fixtures are synthetic but modeled on realistic protocol-drift exports. Clinical and laboratory data managers would use this harmonization to create reproducible, auditable cohorts for downstream scientific analysis and reporting.

Reference ApproachThe intended solution strategy+

A complete solution first infers the uniquely supported decoder for every batch and classifies deliveries by cutoff and supersession. It applies the winning amendment independently per field, constructs all time-valid canonical-subject candidates, and matches them to protocol slots. It then enumerates each group's feasible atomic assignments with exact sum, order, ratio, measurement, and payload rules; searches globally across those assignments using all five objective levels; and derives the cohort, exclusions, deferred groups, lineage, and reconciliation summary from the winning selection.

Verification DesignHow the produced artifacts are independently checked+

The verifier hash-pins the complete shipped contract and inputs, rejects symlinked or non-regular artifacts, strictly parses CSV and finite JSON, and compares every shipped output to protected expected results. It separately copies a different fixture to a read-only directory, blocks access to the expected-data tree, executes the submitted program as uid 65534, and checks all five held-out artifacts. The hidden topology distinguishes a correct global optimizer from locally optimal candidate or priority-greedy implementations.

Task 12 · End-to-end software build

Arcshift Toolpaths

Agentic software engineering · Hardware Embedded and Low Level Systems · CAD and mechanical workflows

Generate safe four-axis CNC programs and execution traces for a custom controller.

Full brief · 6 source files · 2,294 lines · 7 expert hours

Task BriefThe complete agent-facing instruction+

Complete the six ArcShift-4X CNC production jobs listed in /app/spec/work_order.json. The controller behavior and output formats are defined by the normative contract in /app/spec/controller.md. Each work-order record names the machine profile, neutral CAM job, G-code destination, and JSON trace destination for one deliverable.

Create all of these files:

  • /app/output/basic.nc
  • /app/output/basic.trace.json
  • /app/output/indexed.nc
  • /app/output/indexed.trace.json
  • /app/output/quadrants.nc
  • /app/output/quadrants.trace.json
  • /app/output/wrap.nc
  • /app/output/wrap.trace.json
  • /app/output/transitions.nc
  • /app/output/transitions.trace.json
  • /app/output/contours.nc
  • /app/output/contours.trace.json

Every G-code program must reproduce its job's requested tool motion in machine coordinates and obey the corresponding profile's ArcShift syntax, modal state, physical-axis limits, fixture clearances, tool-change position, zone changes, tool-length compensation, rotary epochs, collision-avoidance parking, rapid safety, indexed arc plane and handedness, whole-job rotary costs, fixture-memory lattice gates and ranking, adaptive contour linearization, chained contour acceleration envelopes, curvature-limited feeds, center offsets, and effective-feed rules. Preserve every semantically necessary operation in the job. The extra safe and modal commands explicitly permitted by /app/spec/controller.md are acceptable, but do not add cutting motions or controller-state operations beyond required contour segments, zone selections, tool changes, epoch changes, memory commits, or rotary moves.

Every trace must contain exactly one record per neutral operation, in input order, with the exact keys and state meanings defined in /app/spec/controller.md. Derive each output from its work-order machine profile and job; rules or values from one record must not be carried into another.

Implementation & Verification Code6 authored source files: solution, public tests, and build environment+
Source Browsersolution/solve.py
6 files · 2,294 lines
#!/usr/bin/env python3
"""Reference ArcShift work-order processor."""

from __future__ import annotations

import json
import math
from pathlib import Path


def format_number(value: float) -> str:
    """Format one controller number without negative zero."""
    if abs(value) < 0.0000005:
        value = 0.0
    text = f"{value:.6f}".rstrip("0").rstrip(".")
    return "0" if text in {"", "-0"} else text


def transform(point: list[float], origin: list[float], angle: float) -> list[float]:
    """Rotate a part-local point about X and translate it by the work origin."""
    radians = math.radians(angle)
    x, y, z = map(float, point)
    ox, oy, oz = map(float, origin)
    return [
        ox + x,
        oy + y * math.cos(radians) - z * math.sin(radians),
        oz + y * math.sin(radians) + z * math.cos(radians),
    ]


def distance(left: list[float], right: list[float]) -> float:
    """Return Euclidean distance between two three-dimensional points."""
    return math.sqrt(sum((a - b) ** 2 for a, b in zip(left, right)))


def cubic_point(points: list[list[float]], t: float) -> list[float]:
    """Evaluate a cubic Bezier at one parameter."""
    p0, p1, p2, p3 = points
    u = 1.0 - t
    return [
        u**3 * p0[axis]
        + 3.0 * u**2 * t * p1[axis]
        + 3.0 * u * t**2 * p2[axis]
        + t**3 * p3[axis]
        for axis in range(3)
    ]


def point_segment_distance(
    point: list[float], start: list[float], end: list[float]
) -> float:
    """Measure from a point to the finite chord joining start and end."""
    chord = [end[axis] - start[axis] for axis in range(3)]
    length_squared = sum(value * value for value in chord)
    if length_squared == 0.0:
        return distance(point, start)
    projection = sum(
        (point[axis] - start[axis]) * chord[axis] for axis in range(3)
    ) / length_squared
    projection = max(0.0, min(1.0, projection))
    closest = [start[axis] + projection * chord[axis] for axis in range(3)]
    return distance(point, closest)


def cubic_curvature(points: list[list[float]], t: float) -> float:
    """Calculate the magnitude of spatial cubic curvature at t."""
    p0, p1, p2, p3 = points
    u = 1.0 - t
    first = [
        3.0 * u**2 * (p1[axis] - p0[axis])
        + 6.0 * u * t * (p2[axis] - p1[axis])
        + 3.0 * t**2 * (p3[axis] - p2[axis])
        for axis in range(3)
    ]
    second = [
        6.0 * u * (p2[axis] - 2.0 * p1[axis] + p0[axis])
        + 6.0 * t * (p3[axis] - 2.0 * p2[axis] + p1[axis])
        for axis in range(3)
    ]
    cross = [
        first[1] * second[2] - first[2] * second[1],
        first[2] * second[0] - first[0] * second[2],
        first[0] * second[1] - first[1] * second[0],
    ]
    first_norm = math.sqrt(sum(value * value for value in first))
    cross_norm = math.sqrt(sum(value * value for value in cross))
    return 0.0 if first_norm == 0.0 else cross_norm / first_norm**3


def contour_segments(
    start: list[float],
    operation: dict,
    acceleration: float,
    base_feed: float,
) -> list[tuple[list[float], float]]:
    """Subdivide one cubic and return accepted local endpoints and feeds."""
    points = [
        list(map(float, start)),
        *[list(map(float, point)) for point in operation["control"]],
    ]
    tolerance = float(operation["chord_tolerance"])
    max_length = float(operation["max_segment_length"])
    accepted: list[tuple[list[float], float]] = []

    def visit(t0: float, t1: float, depth: int) -> None:
        assert depth < 20
        begin = cubic_point(points, t0)
        end = cubic_point(points, t1)
        span = t1 - t0
        samples = [
            cubic_point(points, t0 + span * fraction)
            for fraction in (0.25, 0.5, 0.75)
        ]
        deviation = max(
            point_segment_distance(sample, begin, end) for sample in samples
        )
        if deviation <= tolerance and distance(begin, end) <= max_length:
            midpoint = (t0 + t1) / 2.0
            curvature = cubic_curvature(points, midpoint)
            dynamic_cap = (
                math.inf
                if curvature == 0.0
                else 60.0 * math.sqrt(acceleration / curvature)
            )
            accepted.append((end, min(base_feed, dynamic_cap)))
            return
        midpoint = (t0 + t1) / 2.0
        visit(t0, midpoint, depth + 1)
        visit(midpoint, t1, depth + 1)

    visit(0.0, 1.0, 0)
    return accepted


def decompose_angle(logical_a: float) -> tuple[int, float]:
    """Return the ArcShift wrap epoch and half-open physical coordinate."""
    epoch = math.floor((logical_a + 180.0) / 360.0)
    physical = logical_a - 360.0 * epoch
    return epoch, 0.0 if abs(physical) < 0.0000005 else physical


def rotary_park_required(machine: dict, old_a: float, new_a: float) -> bool:
    """Apply the controller's epoch-or-threshold collision-park rule."""
    park = machine.get("rotary_park")
    if park is None:
        return False
    old_epoch, _ = decompose_angle(old_a)
    new_epoch, _ = decompose_angle(new_a)
    return (
        old_epoch != new_epoch
        or abs(new_a - old_a) > float(park["threshold_degrees"])
    )


def plan_rotary_angles(machine: dict, job: dict) -> dict[int, float]:
    """Globally choose windowed index epochs by dynamic programming."""
    indexed = [
        (index, operation)
        for index, operation in enumerate(job["operations"])
        if operation["op"] in {"index", "index_window"}
    ]
    if not any(operation["op"] == "index_window" for _, operation in indexed):
        return {
            index: float(operation["a"]) for index, operation in indexed
        }

    optimizer = machine["rotary_optimizer"]
    assert machine.get("rotary_park") is not None
    penalty = float(optimizer["park_penalty_degrees"])

    # Each state is current logical angle -> (cost, epoch tuple, angle tuple).
    states: dict[float, tuple[float, tuple[int, ...], tuple[float, ...]]] = {
        0.0: (0.0, (), ())
    }
    for _, operation in indexed:
        if operation["op"] == "index":
            candidates = [float(operation["a"])]
        else:
            base = float(operation["a"])
            candidates = [
                base + 360.0 * epoch
                for epoch in range(
                    int(operation["epoch_min"]),
                    int(operation["epoch_max"]) + 1,
                )
            ]

        next_states: dict[
            float, tuple[float, tuple[int, ...], tuple[float, ...]]
        ] = {}
        for candidate in candidates:
            candidate_epoch, _ = decompose_angle(candidate)
            for old_angle, (old_cost, old_epochs, old_angles) in states.items():
                transition = abs(candidate - old_angle)
                if rotary_park_required(machine, old_angle, candidate):
                    transition += penalty
                proposed = (
                    old_cost + transition,
                    old_epochs + (candidate_epoch,),
                    old_angles + (candidate,),
                )
                incumbent = next_states.get(candidate)
                if (
                    incumbent is None
                    or proposed[0] < incumbent[0] - 1e-9
                    or (
                        abs(proposed[0] - incumbent[0]) <= 1e-9
                        and proposed[1] < incumbent[1]
                    )
                ):
                    next_states[candidate] = proposed
        states = next_states

    best = None
    for state in states.values():
        if (
            best is None
            or state[0] < best[0] - 1e-9
            or (
                abs(state[0] - best[0]) <= 1e-9
                and state[1] < best[1]
            )
        ):
            best = state
    assert best is not None
    return {
        operation_index: angle
        for (operation_index, _), angle in zip(indexed, best[2])
    }


def initial_memory_state(machine: dict) -> dict:
    """Return the synthetic fixture-memory state at program start."""
    memory = machine["fixture_memory"]
    return {
        "phase": int(memory["initial_phase"]),
        "debt": int(memory["initial_debt"]),
        "direction": 0,
        "logical": 0.0,
        "epoch": 0,
    }


def apply_memory_pulse(machine: dict, state: dict, operation: dict) -> dict:
    """Apply the per-operation ArcShift pulse before its specific effect."""
    updated = state.copy()
    modulus = int(machine["fixture_memory"]["phase_modulus"])
    updated["phase"] = (
        int(state["phase"])
        + int(operation["memory_pulse"])
        + int(state["debt"])
    ) % modulus
    return updated


def apply_lattice_candidate(machine: dict, state: dict, candidate: dict) -> dict:
    """Advance phase, debt, direction, and rotary state for one lattice slot."""
    updated = state.copy()
    old_phase = int(state["phase"])
    new_epoch = int(candidate["epoch"])
    physical = float(candidate["physical_a"])
    new_logical = physical + 360.0 * new_epoch
    delta = new_logical - float(state["logical"])
    direction = 1 if delta > 0.0 else -1 if delta < 0.0 else 0
    turns = abs(new_epoch - int(state["epoch"]))
    reversal = int(
        int(state["direction"]) != 0
        and direction != 0
        and direction != int(state["direction"])
    )
    modulus = int(machine["fixture_memory"]["phase_modulus"])
    updated["phase"] = (
        old_phase
        + int(candidate["tooth"])
        + direction * (turns + 1)
    ) % modulus
    updated["debt"] = max(
        0,
        int(state["debt"])
        + ((old_phase + int(candidate["slot"])) % 4)
        + turns
        + 2 * reversal
        - int(candidate["relief"]),
    )
    if direction != 0:
        updated["direction"] = direction
    updated["logical"] = new_logical
    updated["epoch"] = new_epoch
    return updated


def memory_gate_accepts(state: dict, operation: dict) -> bool:
    """Return whether the pulsed state satisfies one delayed memory gate."""
    return (
        int(state["phase"]) == int(operation["accept_phase"])
        and int(operation["debt_min"])
        <= int(state["debt"])
        <= int(operation["debt_max"])
    )


def fold_memory_gate(machine: dict, state: dict, operation: dict) -> dict:
    """Fold a committed gate state into the next synthetic memory block."""
    updated = state.copy()
    modulus = int(machine["fixture_memory"]["phase_modulus"])
    updated["phase"] = (
        int(operation["fold_multiplier"]) * int(state["phase"])
        + int(operation["fold_addend"])
    ) % modulus
    updated["debt"] = max(
        0, int(state["debt"]) - int(operation["relief"])
    )
    if bool(operation["flip_direction"]):
        updated["direction"] = -int(state["direction"])
    return updated


def memory_record_better(left: dict, right: dict | None) -> bool:
    """Compare partial or complete fixture-memory paths by contract ranking."""
    if right is None:
        return True
    for key in ("resonance_error", "weighted_debt"):
        if int(left[key]) != int(right[key]):
            return int(left[key]) < int(right[key])
    if float(left["travel_cost"]) < float(right["travel_cost"]) - 1e-9:
        return True
    if abs(float(left["travel_cost"]) - float(right["travel_cost"])) > 1e-9:
        return False
    if tuple(left["slots"]) != tuple(right["slots"]):
        return tuple(left["slots"]) < tuple(right["slots"])
    return tuple(left["epochs"]) < tuple(right["epochs"])


def plan_fixture_memory(machine: dict, job: dict) -> dict:
    """Resolve a memory-enabled job with dynamic programming over ledger state."""
    initial = initial_memory_state(machine)
    states = {
        (
            initial["logical"],
            initial["epoch"],
            initial["phase"],
            initial["debt"],
            initial["direction"],
            0,
        ): {
            "state": initial,
            "peak_debt": 0,
            "resonance_error": int(
                machine["fixture_memory"]["resonance_debt"]
            ),
            "weighted_debt": 0,
            "travel_cost": 0.0,
            "slots": (),
            "epochs": (),
            "lattice_count": 0,
            "angles": {},
            "candidates": {},
            "gates": {},
        }
    }
    debt_limit = int(machine["fixture_memory"]["debt_limit"])
    penalty = float(machine["rotary_optimizer"]["park_penalty_degrees"])

    for operation_index, operation in enumerate(job["operations"]):
        next_states = {}
        for record in states.values():
            pulsed = apply_memory_pulse(machine, record["state"], operation)
            alternatives = []

            if operation["op"] == "index_lattice":
                ordinal = int(record["lattice_count"]) + 1
                for candidate in operation["candidates"]:
                    lower, upper = map(float, machine["physical_a_range"])
                    physical = float(candidate["physical_a"])
                    assert lower <= physical < upper
                    advanced = apply_lattice_candidate(
                        machine, pulsed, candidate
                    )
                    if int(advanced["debt"]) > debt_limit:
                        continue
                    transition = abs(
                        float(advanced["logical"])
                        - float(record["state"]["logical"])
                    )
                    if rotary_park_required(
                        machine,
                        float(record["state"]["logical"]),
                        float(advanced["logical"]),
                    ):
                        transition += penalty
                    proposed = {
                        **record,
                        "state": advanced,
                        "peak_debt": max(
                            int(record["peak_debt"]),
                            int(advanced["debt"]),
                        ),
                        "weighted_debt": int(record["weighted_debt"])
                        + ordinal * int(advanced["debt"]),
                        "travel_cost": float(record["travel_cost"])
                        + transition,
                        "slots": tuple(record["slots"])
                        + (int(candidate["slot"]),),
                        "epochs": tuple(record["epochs"])
                        + (int(candidate["epoch"]),),
                        "lattice_count": ordinal,
                        "angles": {
                            **record["angles"],
                            operation_index: float(advanced["logical"]),
                        },
                        "candidates": {
                            **record["candidates"],
                            operation_index: candidate,
                        },
                    }
                    proposed["resonance_error"] = abs(
                        int(proposed["peak_debt"])
                        - int(
                            machine["fixture_memory"]["resonance_debt"]
                        )
                    )
                    alternatives.append(proposed)
            elif operation["op"] == "memory_gate":
                if not memory_gate_accepts(pulsed, operation):
                    continue
                committed = {
                    "phase": int(pulsed["phase"]),
                    "debt": int(pulsed["debt"]),
                }
                alternatives.append(
                    {
                        **record,
                        "state": fold_memory_gate(
                            machine, pulsed, operation
                        ),
                        "gates": {
                            **record["gates"],
                            operation_index: committed,
                        },
                    }
                )
            else:
                alternatives.append({**record, "state": pulsed})

            for proposed in alternatives:
                state = proposed["state"]
                key = (
                    float(state["logical"]),
                    int(state["epoch"]),
                    int(state["phase"]),
                    int(state["debt"]),
                    int(state["direction"]),
                    int(proposed["peak_debt"]),
                )
                incumbent = next_states.get(key)
                if memory_record_better(proposed, incumbent):
                    next_states[key] = proposed
        states = next_states
        assert states, f"fixture-memory plan has no path after operation {operation_index}"

    best = None
    for record in states.values():
        if memory_record_better(record, best):
            best = record
    assert best is not None
    return {
        "angles": best["angles"],
        "candidates": best["candidates"],
        "gates": best["gates"],
    }


def plan_controller(machine: dict, job: dict) -> dict:
    """Return rotary selections and optional fixture-memory gate states."""
    if "fixture_memory" in machine:
        return plan_fixture_memory(machine, job)
    return {
        "angles": plan_rotary_angles(machine, job),
        "candidates": {},
        "gates": {},
    }


class Postprocessor:
    """Generate one ArcShift controller program and normalized operation trace."""

    def __init__(self, machine: dict, job: dict):
        self.machine = machine
        self.job = job
        self.lines = [f"(PROGRAM {job['program_id']})", "G21", "G90", "G17"]
        self.position = list(map(float, machine["tool_change_position"]))
        self.logical = 0.0
        self.physical = 0.0
        self.epoch = 0
        self.tool = None
        self.zone = None
        self.compensation = False
        self.plane = "G17"
        self.local_point = None
        self.trace = []
        self.contour_plan = {}
        self.memory_state = (
            initial_memory_state(machine)
            if "fixture_memory" in machine
            else None
        )

    def emit(self, line: str) -> None:
        self.lines.append(line)

    def zone_clearance(self, zone: str | None = None) -> float:
        selected = self.zone if zone is None else zone
        if selected is None:
            return float(self.machine["global_clearance_z"])
        return float(self.machine["zones"][selected]["clearance_z"])

    def rotary_clearance(self) -> float:
        return max(
            float(self.machine["global_clearance_z"]),
            float(self.machine["rotary_clearance_z"]),
            self.zone_clearance(),
        )

    def cancel_compensation(self) -> None:
        if self.compensation:
            self.emit("G49")
            self.compensation = False

    def activate_compensation(self) -> None:
        if not self.compensation:
            offset = int(self.machine["tools"][self.tool]["offset"])
            self.emit(f"G43 H{offset}")
            self.compensation = True

    def move_z(self, z: float) -> None:
        if not math.isclose(self.position[2], z, abs_tol=1e-9):
            self.emit(f"G0 Z{format_number(z)}")
            self.position[2] = z

    def retract(self, z: float) -> None:
        if self.position[2] < z:
            self.move_z(z)

    def rapid_machine(self, target: list[float]) -> None:
        target = list(map(float, target))
        x_changed = not math.isclose(self.position[0], target[0], abs_tol=1e-9)
        y_changed = not math.isclose(self.position[1], target[1], abs_tol=1e-9)
        if x_changed or y_changed:
            self.retract(self.zone_clearance())
            self.emit(
                f"G0 X{format_number(target[0])} Y{format_number(target[1])}"
            )
            self.position[0] = target[0]
            self.position[1] = target[1]
        self.move_z(target[2])

    def move_to_tool_change_position(self) -> None:
        target = list(map(float, self.machine["tool_change_position"]))
        self.retract(max(self.zone_clearance(), target[2]))
        x_changed = not math.isclose(self.position[0], target[0], abs_tol=1e-9)
        y_changed = not math.isclose(self.position[1], target[1], abs_tol=1e-9)
        if x_changed or y_changed:
            self.emit(
                f"G0 X{format_number(target[0])} Y{format_number(target[1])}"
            )
            self.position[0] = target[0]
            self.position[1] = target[1]
        self.move_z(target[2])

    def effective_feed(self, operation: dict) -> float:
        base = min(
            float(operation["feed"]),
            float(self.machine["max_linear_feed"]),
            float(self.machine["tools"][self.tool]["max_feed"]),
            float(self.machine["zones"][self.zone]["max_feed"]),
        )
        return base * float(
            self.machine["surface_feed_factors"][operation["surface"]]
        )

    def prepare_contour_chain(self, start_index: int) -> None:
        """Plan one maximal contour chain with forward/backward speed passes."""
        assert self.local_point is not None
        cursor = start_index
        local_start = self.local_point.copy()
        flattened = []

        while (
            cursor < len(self.job["operations"])
            and self.job["operations"][cursor]["op"] == "contour"
        ):
            operation = self.job["operations"][cursor]
            raw_segments = contour_segments(
                local_start,
                operation,
                float(self.machine["max_contour_acceleration"]),
                self.effective_feed(operation),
            )
            previous = local_start
            for local_endpoint, raw_cap in raw_segments:
                flattened.append(
                    {
                        "operation_index": cursor,
                        "local_endpoint": local_endpoint,
                        "length": distance(previous, local_endpoint),
                        "raw_cap": raw_cap,
                    }
                )
                previous = local_endpoint
            local_start = list(map(float, operation["control"][2]))
            cursor += 1

        acceleration = float(
            self.machine["max_contour_tangential_acceleration"]
        )
        node_speeds = [
            float(self.machine["contour_chain_entry_feed"]) / 60.0
        ]
        for segment in flattened:
            reachable = math.sqrt(
                node_speeds[-1] ** 2
                + 2.0 * acceleration * segment["length"]
            )
            node_speeds.append(
                min(float(segment["raw_cap"]) / 60.0, reachable)
            )

        node_speeds[-1] = min(
            node_speeds[-1],
            float(self.machine["contour_chain_exit_feed"]) / 60.0,
        )
        for node in range(len(flattened) - 1, -1, -1):
            reachable = math.sqrt(
                node_speeds[node + 1] ** 2
                + 2.0 * acceleration * flattened[node]["length"]
            )
            node_speeds[node] = min(node_speeds[node], reachable)

        for segment_index, segment in enumerate(flattened, start=1):
            self.contour_plan.setdefault(
                segment["operation_index"], []
            ).append(
                (
                    segment["local_endpoint"],
                    node_speeds[segment_index] * 60.0,
                )
            )

    def record(
        self,
        index: int,
        operation: dict,
        endpoint: list[float] | None = None,
        feed: float | None = None,
    ) -> None:
        self.trace.append(
            {
                "operation_index": index,
                "operation": operation["op"],
                "logical_a": self.logical,
                "physical_a": self.physical,
                "wrap_epoch": self.epoch,
                "machine_endpoint": None if endpoint is None else endpoint.copy(),
                "effective_feed": feed,
            }
        )

    def process(self) -> tuple[list[str], list[dict]]:
        controller_plan = plan_controller(self.machine, self.job)
        planned_angles = controller_plan["angles"]
        for index, operation in enumerate(self.job["operations"]):
            kind = operation["op"]
            endpoint = None
            feed = None
            if self.memory_state is not None:
                self.memory_state = apply_memory_pulse(
                    self.machine, self.memory_state, operation
                )

            if kind == "select_zone":
                new_zone = operation["zone"]
                self.cancel_compensation()
                self.retract(
                    max(self.zone_clearance(), self.zone_clearance(new_zone))
                )
                code = int(self.machine["zones"][new_zone]["code"])
                self.emit(f"M150 P{code}")
                self.zone = new_zone
            elif kind == "tool_change":
                self.cancel_compensation()
                self.move_to_tool_change_position()
                self.tool = operation["tool"]
                self.emit(f"{self.tool} M6")
            elif kind in {"index", "index_window", "index_lattice"}:
                self.cancel_compensation()
                if kind == "index_lattice":
                    self.memory_state = apply_lattice_candidate(
                        self.machine,
                        self.memory_state,
                        controller_plan["candidates"][index],
                    )
                next_logical = planned_angles[index]
                next_epoch, next_physical = decompose_angle(next_logical)
                park = self.machine.get("rotary_park")
                if park is not None and rotary_park_required(
                    self.machine, self.logical, next_logical
                ):
                    self.rapid_machine(park["point"])
                self.retract(self.rotary_clearance())
                self.logical = next_logical
                self.physical = next_physical
                if next_epoch != self.epoch:
                    self.emit(f"M271 P{next_epoch}")
                self.epoch = next_epoch
                self.emit(f"G0 A{format_number(self.physical)}")
                self.local_point = None
            elif kind == "memory_gate":
                expected = controller_plan["gates"][index]
                assert int(self.memory_state["phase"]) == int(
                    expected["phase"]
                )
                assert int(self.memory_state["debt"]) == int(
                    expected["debt"]
                )
                self.cancel_compensation()
                self.rapid_machine(
                    self.machine["fixture_memory"]["commit_position"]
                )
                self.emit(
                    f"M286 P{int(expected['phase'])} Q{int(expected['debt'])}"
                )
                self.memory_state = fold_memory_gate(
                    self.machine, self.memory_state, operation
                )
                self.local_point = None
            elif kind == "rapid":
                endpoint = transform(
                    operation["point"], self.job["origin"], self.logical
                )
                self.rapid_machine(endpoint)
                self.local_point = list(map(float, operation["point"]))
            elif kind == "linear":
                self.activate_compensation()
                endpoint = transform(
                    operation["point"], self.job["origin"], self.logical
                )
                feed = self.effective_feed(operation)
                self.emit(
                    " ".join(
                        [
                            "G1",
                            f"X{format_number(endpoint[0])}",
                            f"Y{format_number(endpoint[1])}",
                            f"Z{format_number(endpoint[2])}",
                            f"F{format_number(feed)}",
                        ]
                    )
                )
                self.position = endpoint.copy()
                self.local_point = list(map(float, operation["point"]))
            elif kind == "arc":
                self.activate_compensation()
                endpoint = transform(
                    operation["end"], self.job["origin"], self.logical
                )
                feed = self.effective_feed(operation)
                q = int(round(self.logical)) % 360
                dx = float(operation["center"][0]) - self.local_point[0]
                dy = float(operation["center"][1]) - self.local_point[1]
                plane, center, reverse = {
                    0: ("G17", {"I": dx, "J": dy}, False),
                    90: ("G18", {"I": dx, "K": dy}, True),
                    180: ("G17", {"I": dx, "J": -dy}, True),
                    270: ("G18", {"I": dx, "K": -dy}, False),
                }[q]
                if self.plane != plane:
                    self.emit(plane)
                    self.plane = plane
                direction = operation["direction"]
                if reverse:
                    direction = "ccw" if direction == "cw" else "cw"
                code = "G3" if direction == "ccw" else "G2"
                center_words = [
                    f"{key}{format_number(value)}" for key, value in center.items()
                ]
                self.emit(
                    " ".join(
                        [
                            code,
                            f"X{format_number(endpoint[0])}",
                            f"Y{format_number(endpoint[1])}",
                            f"Z{format_number(endpoint[2])}",
                            *center_words,
                            f"F{format_number(feed)}",
                        ]
                    )
                )
                self.position = endpoint.copy()
                self.local_point = list(map(float, operation["end"]))
            elif kind == "contour":
                assert self.local_point is not None
                self.activate_compensation()
                if index not in self.contour_plan:
                    self.prepare_contour_chain(index)
                segments = self.contour_plan[index]
                feeds = []
                for local_endpoint, segment_feed in segments:
                    endpoint = transform(
                        local_endpoint, self.job["origin"], self.logical
                    )
                    self.emit(
                        " ".join(
                            [
                                "G1",
                                f"X{format_number(endpoint[0])}",
                                f"Y{format_number(endpoint[1])}",
                                f"Z{format_number(endpoint[2])}",
                                f"F{format_number(segment_feed)}",
                            ]
                        )
                    )
                    self.position = endpoint.copy()
                    feeds.append(segment_feed)
                self.local_point = list(map(float, operation["control"][2]))
                endpoint = self.position.copy()
                feed = min(feeds)
            else:
                raise ValueError(f"unsupported operation {kind}")

            self.record(index, operation, endpoint, feed)

        self.emit("M30")
        return self.lines, self.trace


def main() -> None:
    work_order = json.loads(
        Path("/app/spec/work_order.json").read_text(encoding="utf-8")
    )
    for deliverable in work_order["deliverables"]:
        machine = json.loads(
            Path(deliverable["machine"]).read_text(encoding="utf-8")
        )
        job = json.loads(Path(deliverable["job"]).read_text(encoding="utf-8"))
        lines, trace = Postprocessor(machine, job).process()
        Path(deliverable["gcode"]).write_text(
            "\n".join(lines) + "\n", encoding="utf-8"
        )
        Path(deliverable["trace"]).write_text(
            json.dumps(trace, indent=2) + "\n", encoding="utf-8"
        )


if __name__ == "__main__":
    main()
Evaluation ContractCategory, tested runtime, effort, and required artifacts+
Public identifierevaluation/arcshift-toolpaths
CategoryHardware Embedded and Low Level Systems
SubcategoryCAD and mechanical workflows
Expert estimate7 hours
Model testedOpus-4.8
Agent testedTerminus-2
Required artifacts/app/output/basic.nc, /app/output/basic.trace.json, /app/output/indexed.nc, /app/output/indexed.trace.json, /app/output/quadrants.nc, /app/output/quadrants.trace.json, /app/output/wrap.nc, /app/output/wrap.trace.json, /app/output/transitions.nc, /app/output/transitions.trace.json, /app/output/contours.nc, /app/output/contours.trace.json
Difficulty DesignWhy the benchmark discriminates between plausible and correct work+

This synthetic task models work performed by an experienced CAM/postprocessor engineer integrating rotary machining jobs with a deliberately nonstandard controller. Correct deliverables coordinate local-to-machine geometry, indexed-plane handedness, compensation and clearance state, wrap epochs, and interacting feed limits across six production jobs. Its distinctive wrap job uses an invented ArcShift fixture-memory lattice rather than a customary CNC rule: every neutral operation pulses coupled phase/debt state, delayed gates reject earlier rotary choices, gate folds rewrite the state used by later choices, and a custom peak/weighted-debt ranking outranks ordinary shortest travel. The chosen detents then alter physical orientation, collision parking, contour geometry, modal transitions, and emitted memory commits together. Generic wrap planners and familiar backlash conventions produce plausible but incorrect output. Difficulty comes from constructing and preserving the disclosed custom state model across interacting mechanical and numerical layers, not from data volume, time pressure, or obscure recall; all rules and fixtures are deterministic and synthetic.

Reference ApproachThe intended solution strategy+

The reference implementation reads the work order and uses dynamic programming over logical angle, epoch, fixture-memory phase, debt, and last direction. It applies every operation pulse, expands each explicit lattice slot, prunes debt-limit and delayed-gate failures, folds committed gates, and ranks surviving paths by the contract's custom debt and travel tuple. It then processes each machine/job pair while maintaining zone, tool, compensation, logical and physical rotary state, wrap epoch, local point, and optional memory state. Safety helpers retract in machine Z, visit collision or memory-commit positions, cancel compensation before controller-state changes, and restore the matching offset before cutting. Geometry uses the documented X-axis transform. Cubic contours are recursively subdivided; curvature supplies raw caps, and forward/backward passes derive the maximal chained acceleration envelope. Each work-order record produces its requested program and normalized trace.

Verification DesignHow the produced artifacts are independently checked+

The verifier never executes agent-controlled code. It reads the twelve declared artifacts under /app/output, rejects symlinks and non-regular files, and checks protected input digests before parsing the same bytes. For the custom memory job, independent test-side Cartesian enumeration applies every pulse, candidate update, delayed gate, fold, debt constraint, and ranking component; this is algorithmically independent from the reference dynamic program. Contour feeds are independently derived from all-node path-distance bounds rather than forward/backward passes. Separate geometry derives transforms, arcs, cubic subdivision, curvature caps, parks, and gate commit state. A G-code simulator checks syntax and behavior including compensation, tool and zone changes, clearances, physical-A limits, epochs, collision parks, exact M286 phase/debt commits at the dedicated position, cutting-motion counts, endpoints, and feeds. Atomic tests cover all six jobs. Geometry uses absolute tolerance 0.00001, twenty times six-decimal rounding error while rejecting systematic sub-millimeter bias; feeds use 0.001 for chained floating-point computation without accepting an intentional rule difference.

Task 13 · End-to-end software build

Echofront Signals

Agentic software engineering · Games Puzzles and Interactive Simulation · Game AI and Strategy

Implement an exact behavioral-equilibrium solver for a signal-locked delayed-command game.

Full brief · 6 source files · 2,310 lines · 16 expert hours

Task BriefThe complete agent-facing instruction+

Implement a reusable exact-strategy solver for the signal-locked Echofront game defined by the normative specification at /app/data/RULES.md.

Create /app/echofront.py. It must expose solve_scenarios(document: dict) -> dict, accepting an input document in the specified scenario schema and returning the specified result document without mutating its argument. /app/echofront.py and /app/output.json must be regular files with one filesystem link each, not symbolic or hard links. The program must also support:

python3 /app/echofront.py INPUT_JSON OUTPUT_JSON

The command must read the input document, create or replace OUTPUT_JSON, and write the same result that solve_scenarios returns. Strategies must be equilibria of the complete prior-weighted finite-horizon game, including the first-phase private-signal constraints and every revealed continuation. Solving worlds independently, averaging their values, or returning best responses to one selected opposing plan is not sufficient. Any behavioral equilibrium satisfying the normative pure-plan exploitability conditions is valid.

The supplied cases are in /app/data/scenarios.json, with their expected results in /app/data/sample_expected.json. Run your solver on those cases and leave the resulting document at /app/output.json. Fresh scenarios following the same rules will also be evaluated through the public function.

Implementation & Verification Code6 authored source files: solution, public tests, and build environment+
Source Browsersolution/solve.py
6 files · 2,310 lines
#!/usr/bin/env python3
"""Reference solver for the signal-locked Echofront game."""

from __future__ import annotations

import itertools
import json
import sys
from dataclasses import dataclass
from fractions import Fraction
from functools import lru_cache
from pathlib import Path


@dataclass(frozen=True)
class Edge:
    origin: int
    target: int
    delay: int
    power: int


@dataclass(frozen=True)
class Pulse:
    origin: int
    target: int
    delay: int
    power: int


@dataclass(frozen=True)
class State:
    forces: tuple[int, ...]
    max_reserve: int
    min_reserve: int
    pending: tuple[Pulse, ...]


@dataclass(frozen=True)
class Game:
    node_ids: tuple[str, ...]
    weights: tuple[int, ...]
    edges: tuple[Edge, ...]
    jam_cap: int


@dataclass(frozen=True)
class World:
    prior: Fraction
    max_signal: str
    min_signal: str
    state: State


def sign(value: int) -> int:
    return (value > 0) - (value < 0)


def clamp_force(value: int) -> int:
    return max(-3, min(3, value))


def fraction_text(value: Fraction) -> str:
    return f"{value.numerator}/{value.denominator}"


def resolve(state: State) -> State:
    ready = [pulse for pulse in state.pending if pulse.delay == 0]
    if not ready:
        return state
    controllers = tuple(sign(force) for force in state.forces)
    delta = [0] * len(state.forces)
    for pulse in ready:
        delta[pulse.target] += controllers[pulse.origin] * pulse.power
    forces = tuple(
        clamp_force(force + change)
        for force, change in zip(state.forces, delta)
    )
    waiting = tuple(pulse for pulse in state.pending if pulse.delay != 0)
    return State(forces, state.max_reserve, state.min_reserve, waiting)


def parse_scenario(raw: dict) -> tuple[Game, tuple[World, ...], int]:
    nodes = raw["nodes"]
    node_ids = tuple(node["id"] for node in nodes)
    if len(node_ids) != 3 or len(set(node_ids)) != 3:
        raise ValueError("each scenario must have three unique node ids")
    index = {node_id: i for i, node_id in enumerate(node_ids)}
    weights = tuple(int(node["weight"]) for node in nodes)
    edges = tuple(
        Edge(
            index[edge["origin"]],
            index[edge["target"]],
            int(edge["delay"]),
            int(edge["power"]),
        )
        for edge in raw["edges"]
    )
    edge_shape = {(edge.origin, edge.target, edge.power) for edge in edges}
    game = Game(node_ids, weights, edges, int(raw["jam_cap"]))

    worlds = []
    for world in raw["worlds"]:
        if set(world["forces"]) != set(node_ids):
            raise ValueError("world forces must name every shared node exactly once")
        pending = []
        for pulse in world["pending"]:
            parsed = Pulse(
                index[pulse["origin"]],
                index[pulse["target"]],
                int(pulse["delay"]),
                int(pulse["power"]),
            )
            if (parsed.origin, parsed.target, parsed.power) not in edge_shape:
                raise ValueError("pending pulse does not match a shared edge")
            pending.append(parsed)
        pending.sort(key=lambda pulse: (
            pulse.origin,
            pulse.target,
            pulse.delay,
            pulse.power,
        ))
        reserves = world["reserves"]
        state = State(
            tuple(int(world["forces"][node_id]) for node_id in node_ids),
            int(reserves["max"]),
            int(reserves["min"]),
            tuple(pending),
        )
        worlds.append(
            World(
                Fraction(world["prior"]),
                world["max_signal"],
                world["min_signal"],
                resolve(state),
            )
        )
    if not worlds or sum(world.prior for world in worlds) != 1:
        raise ValueError("world priors must sum to one")
    if any(world.prior <= 0 for world in worlds):
        raise ValueError("world priors must be positive")
    return game, tuple(worlds), int(raw["horizon"])


def legal_actions(game: Game, state: State, player: int) -> tuple[str, ...]:
    reserve = state.max_reserve if player == 1 else state.min_reserve
    actions = ["pass"]
    if reserve <= 0:
        return tuple(actions)

    for index, node_id in enumerate(game.node_ids):
        if sign(state.forces[index]) == player and abs(state.forces[index]) < 3:
            actions.append(f"fortify:{node_id}")

    for edge in game.edges:
        if sign(state.forces[edge.origin]) == player:
            actions.append(
                f"launch:{game.node_ids[edge.origin]}>{game.node_ids[edge.target]}"
            )

    pending_origins = {
        pulse.origin
        for pulse in state.pending
        if pulse.delay < game.jam_cap
    }
    for origin in pending_origins:
        if sign(state.forces[origin]) == -player:
            actions.append(f"jam:{game.node_ids[origin]}")

    return tuple(sorted(set(actions)))


def parse_action(game: Game, action: str) -> tuple[str, int | None, int | None]:
    if action == "pass":
        return "pass", None, None
    kind, payload = action.split(":", 1)
    if kind in {"fortify", "jam"}:
        return kind, game.node_ids.index(payload), None
    origin, target = payload.split(">", 1)
    return kind, game.node_ids.index(origin), game.node_ids.index(target)


def transition(game: Game, state: State, max_action: str, min_action: str) -> State:
    forces = list(state.forces)
    max_reserve = state.max_reserve - (max_action != "pass")
    min_reserve = state.min_reserve - (min_action != "pass")
    pending = list(state.pending)
    parsed = (
        (1, parse_action(game, max_action)),
        (-1, parse_action(game, min_action)),
    )

    for player, (kind, origin, _) in parsed:
        if kind == "fortify":
            forces[origin] = clamp_force(forces[origin] + player)

    jammed_origins = {
        origin
        for _, (kind, origin, _) in parsed
        if kind == "jam"
    }
    pending = [
        Pulse(
            pulse.origin,
            pulse.target,
            min(
                game.jam_cap,
                pulse.delay + int(pulse.origin in jammed_origins),
            ),
            pulse.power,
        )
        for pulse in pending
    ]

    edge_lookup = {(edge.origin, edge.target): edge for edge in game.edges}
    for _, (kind, origin, target) in parsed:
        if kind == "launch":
            edge = edge_lookup[(origin, target)]
            pending.append(Pulse(origin, target, edge.delay, edge.power))

    pending = [
        Pulse(pulse.origin, pulse.target, max(0, pulse.delay - 1), pulse.power)
        for pulse in pending
    ]
    pending.sort(key=lambda pulse: (
        pulse.origin,
        pulse.target,
        pulse.delay,
        pulse.power,
    ))
    return resolve(State(tuple(forces), max_reserve, min_reserve, tuple(pending)))


def terminal_value(game: Game, state: State) -> Fraction:
    board = sum(
        weight * force
        for weight, force in zip(game.weights, state.forces)
    )
    return Fraction(board + state.max_reserve - state.min_reserve)


def solve_linear(
    matrix: list[list[Fraction]],
    rhs: list[Fraction],
) -> list[Fraction] | None:
    size = len(rhs)
    if len(matrix) != size or any(len(row) != size for row in matrix):
        raise ValueError("solve_linear requires a square system")
    augmented = [row[:] + [value] for row, value in zip(matrix, rhs)]
    for column in range(size):
        pivot = next(
            (
                row
                for row in range(column, size)
                if augmented[row][column]
            ),
            None,
        )
        if pivot is None:
            return None
        augmented[column], augmented[pivot] = (
            augmented[pivot],
            augmented[column],
        )
        divisor = augmented[column][column]
        augmented[column] = [
            value / divisor
            for value in augmented[column]
        ]
        for row in range(size):
            if row == column or not augmented[row][column]:
                continue
            factor = augmented[row][column]
            augmented[row] = [
                value - factor * pivot_value
                for value, pivot_value in zip(
                    augmented[row],
                    augmented[column],
                )
            ]
    return [augmented[row][-1] for row in range(size)]


def solve_matrix(
    matrix: tuple[tuple[Fraction, ...], ...],
) -> tuple[Fraction, tuple[Fraction, ...], tuple[Fraction, ...]]:
    rows = len(matrix)
    columns = len(matrix[0])
    if not rows or not columns or any(len(row) != columns for row in matrix):
        raise ValueError("matrix game must be nonempty and rectangular")

    row_floor = [min(row) for row in matrix]
    column_ceiling = [
        max(matrix[row][column] for row in range(rows))
        for column in range(columns)
    ]
    lower = max(row_floor)
    upper = min(column_ceiling)
    if lower == upper:
        row = next(index for index, value in enumerate(row_floor) if value == lower)
        column = next(
            index
            for index, value in enumerate(column_ceiling)
            if value == upper
        )
        return (
            lower,
            tuple(Fraction(index == row) for index in range(rows)),
            tuple(Fraction(index == column) for index in range(columns)),
        )

    for support_size in range(2, min(rows, columns) + 1):
        for row_support in itertools.combinations(range(rows), support_size):
            for column_support in itertools.combinations(
                range(columns),
                support_size,
            ):
                max_system = [
                    [
                        matrix[row][column]
                        for row in row_support
                    ] + [Fraction(-1)]
                    for column in column_support
                ]
                max_system.append(
                    [Fraction(1)] * support_size + [Fraction(0)]
                )
                max_solution = solve_linear(
                    max_system,
                    [Fraction(0)] * support_size + [Fraction(1)],
                )
                if max_solution is None:
                    continue

                min_system = [
                    [
                        matrix[row][column]
                        for column in column_support
                    ] + [Fraction(-1)]
                    for row in row_support
                ]
                min_system.append(
                    [Fraction(1)] * support_size + [Fraction(0)]
                )
                min_solution = solve_linear(
                    min_system,
                    [Fraction(0)] * support_size + [Fraction(1)],
                )
                if min_solution is None:
                    continue

                max_probabilities = max_solution[:-1]
                min_probabilities = min_solution[:-1]
                value = max_solution[-1]
                if (
                    min_solution[-1] != value
                    or any(
                        probability <= 0
                        for probability in max_probabilities + min_probabilities
                    )
                ):
                    continue

                p = [Fraction(0)] * rows
                q = [Fraction(0)] * columns
                for row, probability in zip(row_support, max_probabilities):
                    p[row] = probability
                for column, probability in zip(
                    column_support,
                    min_probabilities,
                ):
                    q[column] = probability
                if any(
                    sum(
                        p[row] * matrix[row][column]
                        for row in range(rows)
                    ) < value
                    for column in range(columns)
                ):
                    continue
                if any(
                    sum(
                        matrix[row][column] * q[column]
                        for column in range(columns)
                    ) > value
                    for row in range(rows)
                ):
                    continue
                return value, tuple(p), tuple(q)

    # Complete rational LP fallback for unequal supports and degenerate faces.
    max_vertices = []
    for active in itertools.combinations(range(columns + rows), rows):
        equations = [[Fraction(1)] * rows + [Fraction(0)]]
        rhs = [Fraction(1)]
        for constraint in active:
            if constraint < columns:
                equations.append(
                    [
                        matrix[row][constraint]
                        for row in range(rows)
                    ] + [Fraction(-1)]
                )
            else:
                row = constraint - columns
                equations.append(
                    [
                        Fraction(index == row)
                        for index in range(rows)
                    ] + [Fraction(0)]
                )
            rhs.append(Fraction(0))
        solution = solve_linear(equations, rhs)
        if solution is None:
            continue
        probabilities = solution[:-1]
        value = solution[-1]
        if any(probability < 0 for probability in probabilities):
            continue
        if any(
            sum(
                probabilities[row] * matrix[row][column]
                for row in range(rows)
            ) < value
            for column in range(columns)
        ):
            continue
        max_vertices.append((value, tuple(probabilities)))

    min_vertices = []
    for active in itertools.combinations(range(rows + columns), columns):
        equations = [[Fraction(1)] * columns + [Fraction(0)]]
        rhs = [Fraction(1)]
        for constraint in active:
            if constraint < rows:
                equations.append(
                    list(matrix[constraint]) + [Fraction(-1)]
                )
            else:
                column = constraint - rows
                equations.append(
                    [
                        Fraction(index == column)
                        for index in range(columns)
                    ] + [Fraction(0)]
                )
            rhs.append(Fraction(0))
        solution = solve_linear(equations, rhs)
        if solution is None:
            continue
        probabilities = solution[:-1]
        value = solution[-1]
        if any(probability < 0 for probability in probabilities):
            continue
        if any(
            sum(
                matrix[row][column] * probabilities[column]
                for column in range(columns)
            ) > value
            for row in range(rows)
        ):
            continue
        min_vertices.append((value, tuple(probabilities)))

    if not max_vertices or not min_vertices:
        raise RuntimeError("matrix game has no feasible rational LP vertex")
    max_value, p = max(max_vertices, key=lambda item: item[0])
    min_value, q = min(min_vertices, key=lambda item: item[0])
    if max_value != min_value:
        raise RuntimeError("matrix game primal and dual values disagree")
    return max_value, p, q


def build_full_information_solver(game: Game):
    @lru_cache(maxsize=None)
    def value(state: State, rounds: int) -> Fraction:
        if rounds == 0:
            return terminal_value(game, state)
        max_actions = legal_actions(game, state, 1)
        min_actions = legal_actions(game, state, -1)
        matrix = tuple(
            tuple(
                value(
                    transition(game, state, max_action, min_action),
                    rounds - 1,
                )
                for min_action in min_actions
            )
            for max_action in max_actions
        )
        return solve_matrix(matrix)[0]

    return value


def signal_actions(
    game: Game,
    worlds: tuple[World, ...],
    player: int,
) -> tuple[tuple[str, tuple[str, ...]], ...]:
    signal_field = (
        (lambda world: world.max_signal)
        if player == 1
        else (lambda world: world.min_signal)
    )
    by_signal: dict[str, tuple[str, ...]] = {}
    for world in worlds:
        signal = signal_field(world)
        actions = legal_actions(game, world.state, player)
        if signal in by_signal and by_signal[signal] != actions:
            raise ValueError(
                f"{signal!r} maps to different legal actions for one player"
            )
        by_signal[signal] = actions
    return tuple(sorted(by_signal.items()))


def pure_policies(
    information: tuple[tuple[str, tuple[str, ...]], ...],
) -> tuple[tuple[str, ...], ...]:
    return tuple(itertools.product(*(actions for _, actions in information)))


def analyze_scenario(raw: dict) -> dict:
    game, worlds, horizon = parse_scenario(raw)
    max_information = signal_actions(game, worlds, 1)
    min_information = signal_actions(game, worlds, -1)
    max_signals = tuple(signal for signal, _ in max_information)
    min_signals = tuple(signal for signal, _ in min_information)
    max_signal_index = {
        signal: index
        for index, signal in enumerate(max_signals)
    }
    min_signal_index = {
        signal: index
        for index, signal in enumerate(min_signals)
    }
    max_policies = pure_policies(max_information)
    min_policies = pure_policies(min_information)
    continuation = build_full_information_solver(game)

    world_tables = []
    for world in worlds:
        max_actions = legal_actions(game, world.state, 1)
        min_actions = legal_actions(game, world.state, -1)
        table = {
            (max_action, min_action): continuation(
                transition(game, world.state, max_action, min_action),
                horizon - 1,
            )
            for max_action in max_actions
            for min_action in min_actions
        }
        world_tables.append(table)

    matrix = tuple(
        tuple(
            sum(
                world.prior
                * world_tables[world_index][
                    (
                        max_policy[max_signal_index[world.max_signal]],
                        min_policy[min_signal_index[world.min_signal]],
                    )
                ]
                for world_index, world in enumerate(worlds)
            )
            for min_policy in min_policies
        )
        for max_policy in max_policies
    )
    value, max_plan_mix, min_plan_mix = solve_matrix(matrix)
    return {
        "game": game,
        "worlds": worlds,
        "horizon": horizon,
        "max_information": max_information,
        "min_information": min_information,
        "max_policies": max_policies,
        "min_policies": min_policies,
        "matrix": matrix,
        "value": value,
        "max_plan_mix": max_plan_mix,
        "min_plan_mix": min_plan_mix,
    }


def behavioral_strategy(
    information: tuple[tuple[str, tuple[str, ...]], ...],
    policies: tuple[tuple[str, ...], ...],
    plan_mix: tuple[Fraction, ...],
) -> dict[str, dict[str, str]]:
    result = {}
    for signal_index, (signal, actions) in enumerate(information):
        probabilities = {
            action: sum(
                probability
                for policy, probability in zip(policies, plan_mix)
                if policy[signal_index] == action
            )
            for action in actions
        }
        result[signal] = {
            action: fraction_text(probabilities[action])
            for action in actions
            if probabilities[action] > 0
        }
    return result


def solve_scenarios(document: dict) -> dict:
    results = []
    for raw in document["scenarios"]:
        analysis = analyze_scenario(raw)
        results.append(
            {
                "id": raw["id"],
                "value": fraction_text(analysis["value"]),
                "max_strategy": behavioral_strategy(
                    analysis["max_information"],
                    analysis["max_policies"],
                    analysis["max_plan_mix"],
                ),
                "min_strategy": behavioral_strategy(
                    analysis["min_information"],
                    analysis["min_policies"],
                    analysis["min_plan_mix"],
                ),
            }
        )
    return {"results": results}


def truth_document(document: dict) -> dict:
    truth = {}
    for raw in document["scenarios"]:
        analysis = analyze_scenario(raw)
        truth[raw["id"]] = {
            "value": fraction_text(analysis["value"]),
            "max_information": [
                {
                    "signal": signal,
                    "actions": list(actions),
                }
                for signal, actions in analysis["max_information"]
            ],
            "min_information": [
                {
                    "signal": signal,
                    "actions": list(actions),
                }
                for signal, actions in analysis["min_information"]
            ],
            "max_policies": [
                list(policy)
                for policy in analysis["max_policies"]
            ],
            "min_policies": [
                list(policy)
                for policy in analysis["min_policies"]
            ],
            "matrix": [
                [fraction_text(value) for value in row]
                for row in analysis["matrix"]
            ],
        }
    return truth


def main() -> None:
    if len(sys.argv) != 3:
        raise SystemExit("usage: echofront.py INPUT_JSON OUTPUT_JSON")
    source, destination = map(Path, sys.argv[1:])
    document = json.loads(source.read_text(encoding="utf-8"))
    destination.parent.mkdir(parents=True, exist_ok=True)
    destination.write_text(
        json.dumps(solve_scenarios(document), indent=2) + "\n",
        encoding="utf-8",
    )


if __name__ == "__main__":
    main()
Evaluation ContractCategory, tested runtime, effort, and required artifacts+
Public identifierevaluation/echofront-signals
CategoryGames Puzzles and Interactive Simulation
SubcategoryGame AI and Strategy
Expert estimate16 hours
Model testedOpus-4.8
Agent testedTerminus-2
Required artifacts/app/echofront.py, /app/output.json
Difficulty DesignWhy the benchmark discriminates between plausible and correct work+

The task composes two exact zero-sum equilibrium layers with different information structures. A solver must first perform queue-complete finite-horizon backward induction after public revelation, then construct the ex-ante Bayesian plan game induced by asymmetric private signal partitions and marginalize a plan equilibrium into behavioral strategies. Independently solving worlds and averaging their values is a persuasive but mathematically invalid near-miss because minmax and expectation do not commute under signal-locked actions. The delayed pulse engine still requires ownership-at-execution, simultaneous snapshots, multiplicity-preserving queues, jam caps, and exact degenerate matrix-game handling. Difficulty is architectural and strategic rather than scale or timeout pressure. The synthetic positions model expert game-AI work on imperfect-information policy engines and adversarial regression suites.

Reference ApproachThe intended solution strategy+

Parse each shared topology and resolve every world's ready pulses before deriving signal information sets. Memoize the perfect-information continuation value by complete post-resolution state and remaining horizon, solving each simultaneous matrix exactly over rational numbers. Enumerate pure first-phase plans as mappings from that player's signals to legal actions. Build the outer zero-sum matrix by prior-weighting each world's continuation payoff, solve that matrix with rational primal and dual programs plus a degenerate fallback, and marginalize the resulting plan mixtures into one behavioral action distribution per signal. Emit the exact ex-ante value and canonical positive rational probabilities.

Verification DesignHow the produced artifacts are independently checked+

The verifier copies only the candidate into a root-owned sandbox, makes /tests unreadable to the candidate uid, and checks function, CLI, immutability, and shipped output contracts while rejecting links, duplicate JSON keys, and non-finite constants. Protected truth freezes complete pure-plan payoff matrices, signal-specific legal action sets, and exact values. Candidate behavioral strategies are expanded into product distributions over all pure plans and checked against every opposing plan, accepting alternate equilibria without hidden tie-breaks. An independent evaluator also builds candidate-hash-specific, structurally fresh signal bundles and computes mixed inner and outer equilibria rather than limiting generated coverage to pure saddles. Cases cover asymmetric signals on either side, nonuniform priors, mixed and degenerate supports, initial simultaneous resolution, queue multiplicity, delayed allegiance conversion, input reordering, and signal partitions whose locally optimal world actions conflict.

Task 14 · End-to-end software build

Cadence Rescue

Agentic software engineering · Data Processing and ETL · Media data processing

Recover progressive video and synchronized PCM from redundant, cadence-converted broadcast field captures.

Full brief · 5 source files · 1,351 lines · 7 expert hours

Task BriefThe complete agent-facing instruction+

Restore the redundant BraidCast archive in /app/data into a progressive preservation master. The capture contains interlaced field packets emitted through segment-specific cadence patterns, redundant retransmissions, damaged packet copies, and coupled parity equations for missing video fields and PCM blocks. The normative input semantics, recovery rules, output formats, validation behavior, and provenance schema are defined in /app/data/CONTRACT.md.

Packet CRC validity is exact: after strict base64 decoding, compute the IEEE CRC-32 over the decoded payload bytes alone, format the masked 32-bit result as eight zero-padded lowercase hexadecimal digits, and compare it case-sensitively with the crc32 string. Decoding, decoded length, and this comparison must all succeed for the packet to be valid.

Create a reusable program at /app/cadence_rescue.py with this interface:

python3 /app/cadence_rescue.py INPUT_DIR OUTPUT_DIR

It will be evaluated on other valid BraidCast archives with different dimensions, cadence patterns, segment layouts, corruptions, duplicate packet topology, and full-rank parity systems. Derive all results from the supplied archive.

Run the program on /app/data and create exactly these deliverables under /app/output:

  • /app/output/master.y4m
  • /app/output/master.wav
  • /app/output/provenance.json

The Y4M file must contain the recovered progressive mono frames in manifest segment and frame order. The WAV file must contain the synchronized recovered PCM blocks in the same order. The provenance report must use the exact normative schema and statistics in /app/data/CONTRACT.md. Do not modify any input file.

Implementation & Verification Code5 authored source files: solution, public tests, and build environment+
Source Browsersolution/cadence_rescue.py
5 files · 1,351 lines
#!/usr/bin/env python3
"""Restore a BraidCast capture into deterministic progressive media masters."""

from __future__ import annotations

import argparse
import base64
import binascii
import hashlib
import json
import os
import shutil
import struct
import tempfile
import wave
import zlib
from pathlib import Path


class CaptureError(ValueError):
    pass


def load_json(path: Path):
    try:
        return json.loads(path.read_text(encoding="utf-8"))
    except (OSError, json.JSONDecodeError) as exc:
        raise CaptureError(f"cannot read {path.name}: {exc}") from exc


def load_jsonl(path: Path):
    records = []
    try:
        with path.open("r", encoding="utf-8") as handle:
            for number, line in enumerate(handle, 1):
                if not line.strip():
                    continue
                try:
                    records.append(json.loads(line))
                except json.JSONDecodeError as exc:
                    raise CaptureError(f"{path.name}:{number}: invalid JSON") from exc
    except OSError as exc:
        raise CaptureError(f"cannot read {path.name}: {exc}") from exc
    return records


def decode_payload(record, expected_size: int):
    try:
        payload = base64.b64decode(record["payload_b64"], validate=True)
        declared_crc = record["crc32"]
    except (KeyError, TypeError, ValueError, UnicodeEncodeError, binascii.Error):
        return None
    if (
        len(payload) != expected_size
        or not isinstance(declared_crc, str)
        or len(declared_crc) != 8
        or any(character not in "0123456789abcdef" for character in declared_crc)
        or f"{zlib.crc32(payload) & 0xFFFFFFFF:08x}" != declared_crc
    ):
        return None
    return payload


def xor_into(target: bytearray, source: bytes):
    for index, value in enumerate(source):
        target[index] ^= value


def recover_vectors(expected_keys, known, equations, vector_size: int):
    missing = [key for key in expected_keys if key not in known]
    if not missing:
        for terms, parity in equations:
            check = bytearray(vector_size)
            for key in terms:
                xor_into(check, known[key])
            if bytes(check) != parity:
                raise CaptureError("parity equation contradicts data packets")
        return known, []

    column = {key: index for index, key in enumerate(missing)}
    rows = []
    for terms, parity in equations:
        mask = 0
        rhs = bytearray(parity)
        for key in terms:
            if key not in expected_keys:
                raise CaptureError("parity equation references an unknown media unit")
            if key in known:
                xor_into(rhs, known[key])
            else:
                mask ^= 1 << column[key]
        if mask == 0:
            if any(rhs):
                raise CaptureError("parity equation is inconsistent")
        else:
            rows.append([mask, rhs])

    pivot_row = 0
    pivots = {}
    for col in range(len(missing)):
        selected = next(
            (row_index for row_index in range(pivot_row, len(rows))
             if rows[row_index][0] & (1 << col)),
            None,
        )
        if selected is None:
            continue
        rows[pivot_row], rows[selected] = rows[selected], rows[pivot_row]
        pivot_mask, pivot_rhs = rows[pivot_row]
        for row_index, (mask, rhs) in enumerate(rows):
            if row_index != pivot_row and mask & (1 << col):
                rows[row_index][0] ^= pivot_mask
                xor_into(rows[row_index][1], pivot_rhs)
        pivots[col] = pivot_row
        pivot_row += 1

    if len(pivots) != len(missing):
        raise CaptureError("braided parity does not uniquely recover every missing unit")

    recovered = dict(known)
    for col, key in enumerate(missing):
        row = rows[pivots[col]]
        if row[0] != 1 << col:
            raise CaptureError("braided parity system is not uniquely reduced")
        recovered[key] = bytes(row[1])

    for terms, parity in equations:
        check = bytearray(vector_size)
        for key in terms:
            xor_into(check, recovered[key])
        if bytes(check) != parity:
            raise CaptureError("recovered media does not satisfy every parity equation")
    return recovered, missing


def validate_manifest(manifest):
    if manifest.get("format") != "braidcast-capture-v1":
        raise CaptureError("unsupported capture format")
    video = manifest.get("video")
    audio = manifest.get("audio")
    segments = manifest.get("segments")
    if not isinstance(video, dict) or not isinstance(audio, dict) or not isinstance(segments, list):
        raise CaptureError("manifest is missing video, audio, or segments")
    width = video.get("width")
    height = video.get("height")
    if (
        type(width) is not int
        or type(height) is not int
        or width <= 0
        or height <= 0
        or height % 2
        or video.get("chroma") != "mono8"
        or video.get("frame_rate") != {"numerator": 24000, "denominator": 1001}
    ):
        raise CaptureError("unsupported video contract")
    if (
        audio.get("sample_rate") != 48000
        or audio.get("channels") != 1
        or audio.get("sample_format") != "s16le"
        or audio.get("samples_per_frame") != 2002
    ):
        raise CaptureError("unsupported audio contract")

    segment_map = {}
    for segment in segments:
        if not isinstance(segment, dict):
            raise CaptureError("segment must be an object")
        segment_id = segment.get("id")
        frame_count = segment.get("frame_count")
        pattern = segment.get("cadence_pattern")
        if (
            not isinstance(segment_id, str)
            or not segment_id
            or segment_id in segment_map
            or type(frame_count) is not int
            or frame_count <= 0
            or not isinstance(pattern, list)
            or not pattern
        ):
            raise CaptureError("invalid segment declaration")
        frame_deltas = [
            item.get("frame_delta")
            for item in pattern
            if isinstance(item, dict) and type(item.get("frame_delta")) is int
        ]
        if len(frame_deltas) != len(pattern) or min(frame_deltas) < 0:
            raise CaptureError("invalid cadence pattern")
        cycle_frames = max(frame_deltas) + 1
        if frame_count % cycle_frames:
            raise CaptureError("segment frame count is not cadence-aligned")
        cadence_counts = {}
        for item in pattern:
            if (
                not isinstance(item, dict)
                or type(item.get("frame_delta")) is not int
                or item["frame_delta"] not in range(cycle_frames)
                or item.get("parity") not in ("top", "bottom")
            ):
                raise CaptureError("invalid cadence pattern")
            identity = (item["frame_delta"], item["parity"])
            cadence_counts[identity] = cadence_counts.get(identity, 0) + 1
        expected_identities = {
            (frame, parity)
            for frame in range(cycle_frames)
            for parity in ("top", "bottom")
        }
        if set(cadence_counts) != expected_identities:
            raise CaptureError("cadence pattern does not cover its progressive frames")
        segment_map[segment_id] = segment
    if not segment_map:
        raise CaptureError("capture has no segments")
    return width, height, segment_map


def video_key_from_packet(record, segment_map):
    try:
        segment_id = record["segment_id"]
        ordinal = record["ordinal"]
        capture = record["capture"]
    except KeyError as exc:
        raise CaptureError("video packet is missing identity fields") from exc
    if (
        segment_id not in segment_map
        or type(ordinal) is not int
        or ordinal < 0
        or not isinstance(capture, str)
        or not capture
    ):
        raise CaptureError("invalid video packet identity")
    segment = segment_map[segment_id]
    pattern = segment["cadence_pattern"]
    cycle_frames = max(item["frame_delta"] for item in pattern) + 1
    cycle, slot = divmod(ordinal, len(pattern))
    entry = pattern[slot]
    frame_index = cycle * cycle_frames + entry["frame_delta"]
    if frame_index >= segment["frame_count"]:
        raise CaptureError("video packet ordinal lies outside its segment")
    return segment_id, frame_index, entry["parity"]


def direct_key(record, segment_map, kind: str):
    try:
        segment_id = record["segment_id"]
        index = record["frame_index"] if kind == "video" else record["block_index"]
    except KeyError as exc:
        raise CaptureError(f"{kind} parity term is missing identity fields") from exc
    if segment_id not in segment_map or type(index) is not int:
        raise CaptureError(f"invalid {kind} parity term")
    if index < 0 or index >= segment_map[segment_id]["frame_count"]:
        raise CaptureError(f"{kind} parity term lies outside its segment")
    if kind == "video":
        parity = record.get("parity")
        if parity not in ("top", "bottom"):
            raise CaptureError("invalid video parity term")
        return segment_id, index, parity
    return segment_id, index


def audio_key_from_packet(record, segment_map):
    capture = record.get("capture")
    if not isinstance(capture, str) or not capture:
        raise CaptureError("invalid audio packet identity")
    return direct_key(record, segment_map, "audio")


def collect_data_packets(records, expected_keys, key_function, vector_size: int):
    known = {}
    valid = 0
    rejected = 0
    for record in records:
        key = key_function(record)
        if key not in expected_keys:
            raise CaptureError("data packet maps outside the declared media")
        payload = decode_payload(record, vector_size)
        if payload is None:
            rejected += 1
            continue
        valid += 1
        previous = known.get(key)
        if previous is not None and previous != payload:
            raise CaptureError("valid redundant packets disagree")
        known[key] = payload
    return known, valid, rejected, valid - len(known)


def collect_parity_packets(records, expected_keys, term_function, vector_size: int):
    equations = []
    valid = 0
    rejected = 0
    identifiers = set()
    for record in records:
        identifier = record.get("equation_id")
        terms_raw = record.get("terms")
        if not isinstance(identifier, str) or not identifier or identifier in identifiers:
            raise CaptureError("invalid or duplicate parity equation id")
        identifiers.add(identifier)
        if not isinstance(terms_raw, list) or not terms_raw:
            raise CaptureError("parity equation has no terms")
        terms = [term_function(term) for term in terms_raw]
        if len(set(terms)) != len(terms) or any(term not in expected_keys for term in terms):
            raise CaptureError("parity equation contains invalid or duplicate terms")
        payload = decode_payload(record, vector_size)
        if payload is None:
            rejected += 1
            continue
        valid += 1
        equations.append((terms, payload))
    return equations, valid, rejected


def y4m_bytes(width: int, height: int, frames):
    content = bytearray(
        f"YUV4MPEG2 W{width} H{height} F24000:1001 Ip A0:0 Cmono\n".encode("ascii")
    )
    for frame in frames:
        content.extend(b"FRAME\n")
        content.extend(frame)
    return bytes(content)


def wav_bytes(pcm: bytes):
    buffer_path = tempfile.NamedTemporaryFile(delete=False)
    buffer_path.close()
    try:
        with wave.open(buffer_path.name, "wb") as output:
            output.setnchannels(1)
            output.setsampwidth(2)
            output.setframerate(48000)
            output.writeframes(pcm)
        return Path(buffer_path.name).read_bytes()
    finally:
        try:
            os.unlink(buffer_path.name)
        except FileNotFoundError:
            pass


def unit_name(key, kind: str):
    if kind == "video":
        return f"{key[0]}:{key[1]:06d}:{key[2]}"
    return f"{key[0]}:{key[1]:06d}"


def restore(input_dir: Path):
    manifest = load_json(input_dir / "manifest.json")
    width, height, segment_map = validate_manifest(manifest)
    segments = manifest["segments"]
    field_size = width * (height // 2)
    audio_block_size = manifest["audio"]["samples_per_frame"] * 2

    expected_video = []
    expected_audio = []
    for segment in segments:
        for frame_index in range(segment["frame_count"]):
            expected_video.extend(
                [(segment["id"], frame_index, "top"), (segment["id"], frame_index, "bottom")]
            )
            expected_audio.append((segment["id"], frame_index))
    video_set = set(expected_video)
    audio_set = set(expected_audio)

    video_known, video_valid, video_rejected, video_collapsed = collect_data_packets(
        load_jsonl(input_dir / "video_packets.jsonl"),
        video_set,
        lambda record: video_key_from_packet(record, segment_map),
        field_size,
    )
    video_equations, video_parity_valid, video_parity_rejected = collect_parity_packets(
        load_jsonl(input_dir / "video_parity.jsonl"),
        video_set,
        lambda term: direct_key(term, segment_map, "video"),
        field_size,
    )
    video_values, recovered_video = recover_vectors(
        expected_video, video_known, video_equations, field_size
    )

    audio_known, audio_valid, audio_rejected, audio_collapsed = collect_data_packets(
        load_jsonl(input_dir / "audio_packets.jsonl"),
        audio_set,
        lambda record: audio_key_from_packet(record, segment_map),
        audio_block_size,
    )
    audio_equations, audio_parity_valid, audio_parity_rejected = collect_parity_packets(
        load_jsonl(input_dir / "audio_parity.jsonl"),
        audio_set,
        lambda term: direct_key(term, segment_map, "audio"),
        audio_block_size,
    )
    audio_values, recovered_audio = recover_vectors(
        expected_audio, audio_known, audio_equations, audio_block_size
    )

    frames = []
    pcm = bytearray()
    segment_reports = []
    output_frame = 0
    for segment in segments:
        segment_id = segment["id"]
        segment_pixels = bytearray()
        segment_pcm = bytearray()
        first_output_frame = output_frame
        for frame_index in range(segment["frame_count"]):
            top = video_values[(segment_id, frame_index, "top")]
            bottom = video_values[(segment_id, frame_index, "bottom")]
            frame = bytearray(width * height)
            top_offset = bottom_offset = 0
            for row in range(height):
                if row % 2 == 0:
                    frame[row * width:(row + 1) * width] = top[top_offset:top_offset + width]
                    top_offset += width
                else:
                    frame[row * width:(row + 1) * width] = bottom[
                        bottom_offset:bottom_offset + width
                    ]
                    bottom_offset += width
            frame_bytes = bytes(frame)
            block = audio_values[(segment_id, frame_index)]
            frames.append(frame_bytes)
            pcm.extend(block)
            segment_pixels.extend(frame_bytes)
            segment_pcm.extend(block)
            output_frame += 1
        segment_reports.append(
            {
                "id": segment_id,
                "first_output_frame": first_output_frame,
                "frame_count": segment["frame_count"],
                "video_sha256": hashlib.sha256(segment_pixels).hexdigest(),
                "audio_sha256": hashlib.sha256(segment_pcm).hexdigest(),
            }
        )

    y4m = y4m_bytes(width, height, frames)
    wav = wav_bytes(bytes(pcm))
    report = {
        "format": "braidcast-report-v1",
        "video": {
            "frame_count": len(frames),
            "field_bytes": field_size,
            "valid_data_packets": video_valid,
            "rejected_data_packets": video_rejected,
            "collapsed_duplicates": video_collapsed,
            "valid_parity_packets": video_parity_valid,
            "rejected_parity_packets": video_parity_rejected,
            "recovered_fields": sorted(unit_name(key, "video") for key in recovered_video),
            "sha256": hashlib.sha256(y4m).hexdigest(),
        },
        "audio": {
            "sample_count": len(pcm) // 2,
            "valid_data_packets": audio_valid,
            "rejected_data_packets": audio_rejected,
            "collapsed_duplicates": audio_collapsed,
            "valid_parity_packets": audio_parity_valid,
            "rejected_parity_packets": audio_parity_rejected,
            "recovered_blocks": sorted(unit_name(key, "audio") for key in recovered_audio),
            "sha256": hashlib.sha256(wav).hexdigest(),
        },
        "segments": segment_reports,
    }
    return y4m, wav, json.dumps(report, indent=2, sort_keys=True).encode("utf-8") + b"\n"


def commit_output(output_dir: Path, files):
    parent = output_dir.parent
    parent.mkdir(parents=True, exist_ok=True)
    staging = Path(tempfile.mkdtemp(prefix=f".{output_dir.name}.staging-", dir=parent))
    try:
        (staging / "master.y4m").write_bytes(files[0])
        (staging / "master.wav").write_bytes(files[1])
        (staging / "provenance.json").write_bytes(files[2])
        if output_dir.exists():
            if output_dir.is_symlink() or not output_dir.is_dir():
                raise CaptureError("output path must be a directory")
            shutil.rmtree(output_dir)
        staging.replace(output_dir)
    except Exception:
        shutil.rmtree(staging, ignore_errors=True)
        raise


def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("input_dir")
    parser.add_argument("output_dir")
    args = parser.parse_args()
    try:
        files = restore(Path(args.input_dir))
        commit_output(Path(args.output_dir), files)
    except (CaptureError, OSError, KeyError, TypeError, struct.error) as exc:
        print(f"cadence rescue failed: {exc}", file=os.sys.stderr)
        return 2
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
Evaluation ContractCategory, tested runtime, effort, and required artifacts+
Public identifierevaluation/cadence-rescue
CategoryData Processing and ETL
SubcategoryMedia data processing
Expert estimate7 hours
Model testedOpus-4.8
Agent testedTerminus-2
Required artifacts/app/cadence_rescue.py, /app/output/master.y4m, /app/output/master.wav, /app/output/provenance.json
Difficulty DesignWhy the benchmark discriminates between plausible and correct work+

This task models a preservation engineer recovering a damaged broadcast ingest whose redundant interlaced captures were produced through different segment-local cadence geometries. A correct implementation must derive each cadence cycle from the manifest, map capture ordinals back to progressive field identities, reject bad CRC copies without discarding valid retransmissions, collapse cadence and source duplicates, solve a custom braided GF(2) erasure system over whole raster fields and PCM blocks, verify every parity relation, weave scanlines, preserve segment order and audio sample alignment, emit exact Y4M and RIFF/WAVE containers, and report deterministic provenance. Missing units are deliberately coupled so single-equation repair and greedy packet selection fail while producing plausible media. The fixtures are deterministic synthetic captures built from small but realistic mono raster and signed PCM essence; archive, restoration, and broadcast-quality engineers perform analogous cadence, redundancy, FEC, synchronization, and conformance work when salvaging contribution feeds or legacy mezzanine assets.

Reference ApproachThe intended solution strategy+

Parse and validate the manifest first, then expand every segment into its expected progressive video-field and audio-block identities. Derive each segment's cycle width and frame span from its cadence pattern, map every video packet ordinal through that geometry, validate payload length and CRC, and collapse identical valid copies while rejecting damaged ones. Reduce each valid parity payload by XORing known terms, construct a binary coefficient matrix for the remaining media units, and use Gaussian elimination over GF(2) while applying the same row operations bytewise to each payload vector. Require full rank, recover every missing field or block, and recheck all equations. Interleave top-field rows into even scanlines and bottom-field rows into odd scanlines, concatenate audio blocks in the identical segment/frame order, serialize a mono 24000/1001 Y4M stream and canonical mono 48 kHz PCM16 WAV, derive the declared hashes and packet statistics, and replace the output directory only after the entire archive validates.

Verification DesignHow the produced artifacts are independently checked+

The verifier checks the four declared artifacts and compares the shipped Y4M, WAV, and provenance bytes against expected media assembled directly from protected pristine frame and audio truth. It then executes the submitted program as an unprivileged user on the shipped archive, a verifier-only mixed-cadence archive, and a runtime-retimed variant whose packet ordinals use different two-frame and eight-frame cadence-cycle geometries. Expected media are constructed from protected source essence rather than imported from the submitted program or a replaceable helper module. Positive tests require exact recovery across corruption, duplicates, coupled missing-unit topology, variable cadence slot counts, variable cycle spans, and complete stale-output replacement. A systematic fatal-boundary matrix covers unknown and out-of-range data and parity identities, conflicting CRC-valid copies, duplicate parity terms and equation ids, unresolved systems, contradictory parity, and unsupported format, media, segment, and cadence declarations; every fatal case must exit nonzero without changing an existing output directory. Recoverable strict-base64, decoded-length, lowercase-CRC-format, and CRC-mismatch cases must instead be counted and recovered. All media, packet order, CRCs, parity equations, and expected bytes are deterministic; exact comparison is appropriate because the contract fixes integer XOR recovery, scanline order, sample order, container headers, JSON serialization, and hashes.

Task 15 · End-to-end software build

Weave Orbit Census

Agentic software engineering · Mathematics and Formal Reasoning · Combinatorics and enumeration

Implement an exact orbit census for cyclic multi-stride equality weaves.

Full brief · 5 source files · 1,126 lines · 8 expert hours

Task BriefThe complete agent-facing instruction+

Implement an exact enumerator for the equality-weave instance in /app/data/instance.json. Write a self-contained Python program at /app/weave_orbits.py with this interface:

python3 /app/weave_orbits.py INSTANCE OUTPUT

Both arguments are absolute paths. Run it on the shipped instance and create /app/output.json.

An equality weave is a cyclic word w[0]...w[n-1] over the integers 0 through q-1, where q is alphabet_size and n is length. Indices below are modulo n. A weave is valid exactly when:

  • every symbol occurs, w[i] != w[i+1] for every i, and the sorted symbol-frequency list equals multiplicities;
  • for every decimal-string key d in distance_matches, exactly its associated number of indices satisfy w[i] == w[i+d];
  • every entry in motif_profiles has the exact motif ledger described below.

For a sequence, its restricted-growth code is obtained by renaming the first distinct value encountered to 0, the next previously unseen value to 1, and so on. The code of a motif window is the lexicographically smaller string between the restricted-growth code of the window and that of its reversal. For a profile with width = k and stride = s, form one window for every start i:

w[i], w[i+s], ..., w[i+(k-1)s]

The frequency map of their codes must equal counts exactly; an omitted code has frequency zero.

Every supplied instance is valid: 2 <= q <= 9; n >= max(q, 3); multiplicities is a sorted positive partition of n with q parts; distance keys are distinct values from 2 through n-1; and motif (width, stride) pairs are unique. A motif width is from 3 through n-1, its stride is from 1 through n-1, its sampled positions are distinct, and its positive listed counts sum to n.

For an orientation e in {1, -1}, offset r, and symbol permutation p, define the group action by T(w)[i] = p(w[r + e*i]). Two valid weaves are equivalent when one is T(w) for some such choices. A representative is the lexicographically smallest digit string in its equivalence class.

OUTPUT must be a JSON object with exactly these keys:

  • alphabet_size, length, and group_order, where group_order = 2 * n * q!;
  • valid_labeled_words, the number of valid words on the labeled alphabet;
  • orbit_count;
  • orbit_size_histogram, mapping decimal orbit-size strings to counts, ordered numerically by orbit size;
  • orbits, ordered by representative. Each entry has exactly representative, orbit_size, and stabilizer_size, with orbit_size * stabilizer_size = group_order.
  • burnside_fixed_sum, the sum of all fixed-word counts, equal to group_order * orbit_count;
  • fixed_point_table, one entry for every group element. Order first by orientation (1 before -1), then increasing offset, then lexicographic permutation string. Each entry has exactly orientation, offset, permutation, and fixed_valid_words. The permutation string is p(0)p(1)...p(q-1), and fixed_valid_words counts valid labeled words satisfying T(w) = w.

JSON integers must be exact. Representatives contain exactly n digits. Use only the Python standard library, derive all results from INSTANCE, and leave /app/data/instance.json unchanged.

Implementation & Verification Code5 authored source files: solution, public tests, and build environment+
Source Browsersolution/solve.py
5 files · 1,126 lines
#!/usr/bin/env python3
"""Reference enumerator for multi-stride equality weaves."""

from __future__ import annotations

import argparse
import itertools
import json
import math
from collections import Counter
from pathlib import Path
from typing import Iterable, Sequence


def restricted_growth(values: Sequence[int]) -> tuple[int, ...]:
    """Rename symbols by order of first appearance."""
    names: dict[int, int] = {}
    next_name = 0
    result: list[int] = []
    for value in values:
        if value not in names:
            names[value] = next_name
            next_name += 1
        result.append(names[value])
    return tuple(result)


def motif_code(values: Sequence[int]) -> str:
    """Canonical equality code, identifying a window with its reversal."""
    forward = restricted_growth(values)
    backward = restricted_growth(tuple(reversed(values)))
    return "".join(str(value) for value in min(forward, backward))


def normalized_dihedral_images(word: Sequence[int]) -> set[tuple[int, ...]]:
    """Dihedral images after quotienting out all symbol relabelings."""
    n = len(word)
    images: set[tuple[int, ...]] = set()
    for offset in range(n):
        images.add(restricted_growth(tuple(word[(offset + i) % n] for i in range(n))))
        images.add(restricted_growth(tuple(word[(offset - i) % n] for i in range(n))))
    return images


def canonical_word(word: Sequence[int]) -> tuple[int, ...]:
    return min(normalized_dihedral_images(word))


def canonical_string(word: Sequence[int]) -> str:
    return "".join(str(value) for value in word)


def parse_instance(path: Path) -> dict:
    raw = json.loads(path.read_text(encoding="utf-8"))
    required = {
        "alphabet_size",
        "length",
        "multiplicities",
        "distance_matches",
        "motif_profiles",
    }
    if set(raw) != required:
        raise ValueError(f"instance keys must be exactly {sorted(required)}")

    q = raw["alphabet_size"]
    n = raw["length"]
    multiplicities = raw["multiplicities"]
    if not isinstance(q, int) or not 2 <= q <= 9:
        raise ValueError("alphabet_size must be an integer from 2 through 9")
    if not isinstance(n, int) or n < q or n < 3:
        raise ValueError("length must be an integer at least alphabet_size and 3")
    if (
        not isinstance(multiplicities, list)
        or len(multiplicities) != q
        or any(not isinstance(value, int) or value <= 0 for value in multiplicities)
        or multiplicities != sorted(multiplicities)
        or sum(multiplicities) != n
    ):
        raise ValueError("multiplicities must be a sorted positive partition of length")

    distance_targets: dict[int, int] = {}
    if not isinstance(raw["distance_matches"], dict):
        raise ValueError("distance_matches must be an object")
    for key, value in raw["distance_matches"].items():
        if not isinstance(key, str) or not key.isdigit():
            raise ValueError("distance keys must be decimal strings")
        distance = int(key)
        if not 2 <= distance < n:
            raise ValueError("each distance must be at least 2 and less than length")
        if not isinstance(value, int) or not 0 <= value <= n:
            raise ValueError("distance match counts must lie between 0 and length")
        distance_targets[distance] = value

    profiles: list[tuple[int, int, dict[str, int]]] = []
    if not isinstance(raw["motif_profiles"], list) or not raw["motif_profiles"]:
        raise ValueError("motif_profiles must be a nonempty array")
    seen_profile_keys: set[tuple[int, int]] = set()
    for profile in raw["motif_profiles"]:
        if not isinstance(profile, dict) or set(profile) != {"width", "stride", "counts"}:
            raise ValueError("each motif profile needs exactly width, stride, and counts")
        width = profile["width"]
        stride = profile["stride"]
        counts = profile["counts"]
        if not isinstance(width, int) or not 3 <= width < n:
            raise ValueError("motif width must be at least 3 and less than length")
        if not isinstance(stride, int) or not 1 <= stride < n:
            raise ValueError("motif stride must lie between 1 and length - 1")
        if len({(step * stride) % n for step in range(width)}) != width:
            raise ValueError("a motif profile may not revisit a position")
        if (width, stride) in seen_profile_keys:
            raise ValueError("motif (width, stride) pairs must be unique")
        seen_profile_keys.add((width, stride))
        if not isinstance(counts, dict):
            raise ValueError("motif counts must be an object")
        normalized_counts: dict[str, int] = {}
        for code, count in counts.items():
            if (
                not isinstance(code, str)
                or len(code) != width
                or not code.isdigit()
                or any(int(char) >= q for char in code)
                or tuple(int(char) for char in code)
                != restricted_growth(tuple(int(char) for char in code))
                or code != motif_code(tuple(int(char) for char in code))
            ):
                raise ValueError(f"invalid canonical motif code {code!r}")
            if not isinstance(count, int) or count <= 0:
                raise ValueError("listed motif counts must be positive")
            normalized_counts[code] = count
        if sum(normalized_counts.values()) != n:
            raise ValueError("each motif profile must count exactly length windows")
        profiles.append((width, stride, normalized_counts))

    return {
        "q": q,
        "n": n,
        "multiplicities": tuple(multiplicities),
        "distance_targets": distance_targets,
        "profiles": profiles,
    }


def counts_can_reach_partition(
    counts: Sequence[int], used_symbols: int, q: int, target: Sequence[int]
) -> bool:
    padded = list(counts[:used_symbols]) + [0] * (q - used_symbols)
    return all(current <= capacity for current, capacity in zip(sorted(padded), target))


def enumerate_normalized_words(
    instance: dict,
) -> tuple[list[tuple[int, ...]], set[tuple[int, ...]]]:
    q: int = instance["q"]
    n: int = instance["n"]
    target: tuple[int, ...] = instance["multiplicities"]
    distance_targets: dict[int, int] = instance["distance_targets"]
    profiles: list[tuple[int, int, dict[str, int]]] = instance["profiles"]

    word = [0]
    counts = [1] + [0] * (q - 1)
    nonwrap_matches = {distance: 0 for distance in distance_targets}
    profile_counts = [Counter() for _ in profiles]
    windows_completed_at: list[list[tuple[int, tuple[int, ...]]]] = [
        [] for _ in range(n)
    ]
    for profile_index, (width, stride, _) in enumerate(profiles):
        for start in range(n):
            indices = tuple(
                (start + step * stride) % n for step in range(width)
            )
            windows_completed_at[max(indices)].append(
                (profile_index, indices)
            )
    normalized_words: list[tuple[int, ...]] = []
    canonical_representatives: set[tuple[int, ...]] = set()

    def visit(max_symbol: int) -> None:
        position = len(word)
        if position == n:
            if word[-1] == word[0] or max_symbol != q - 1:
                return
            if tuple(sorted(counts)) != target:
                return
            for distance, target_count in distance_targets.items():
                wrap = sum(
                    word[index] == word[(index + distance) % n]
                    for index in range(n - distance, n)
                )
                if nonwrap_matches[distance] + wrap != target_count:
                    return
            if any(
                dict(profile_counts[index]) != target_counts
                for index, (_, _, target_counts) in enumerate(profiles)
            ):
                return
            completed = tuple(word)
            normalized_words.append(completed)
            canonical_representatives.add(canonical_word(completed))
            return

        remaining_after = n - position - 1
        largest_allowed = min(max_symbol + 1, q - 1)
        for symbol in range(largest_allowed + 1):
            if symbol == word[-1]:
                continue
            new_max = max(max_symbol, symbol)
            unseen_after = q - (new_max + 1)
            if unseen_after > remaining_after:
                continue

            counts[symbol] += 1
            if counts_can_reach_partition(counts, new_max + 1, q, target):
                increments: list[tuple[int, int]] = []
                feasible = True
                for distance, target_count in distance_targets.items():
                    increment = int(position >= distance and word[position - distance] == symbol)
                    nonwrap_matches[distance] += increment
                    increments.append((distance, increment))
                    if nonwrap_matches[distance] > target_count:
                        feasible = False
                if feasible:
                    word.append(symbol)
                    motif_increments: list[tuple[int, str]] = []
                    for profile_index, indices in windows_completed_at[position]:
                        code = motif_code(tuple(word[index] for index in indices))
                        profile_counts[profile_index][code] += 1
                        motif_increments.append((profile_index, code))
                        target_counts = profiles[profile_index][2]
                        if (
                            profile_counts[profile_index][code]
                            > target_counts.get(code, 0)
                        ):
                            feasible = False
                    if feasible:
                        visit(new_max)
                    for profile_index, code in motif_increments:
                        profile_counts[profile_index][code] -= 1
                        if profile_counts[profile_index][code] == 0:
                            del profile_counts[profile_index][code]
                    word.pop()
                for distance, increment in increments:
                    nonwrap_matches[distance] -= increment
            counts[symbol] -= 1

    visit(0)
    return normalized_words, canonical_representatives


def labeled_words(
    normalized_words: Iterable[Sequence[int]], q: int
) -> list[tuple[int, ...]]:
    return [
        tuple(permutation[value] for value in word)
        for word in normalized_words
        for permutation in itertools.permutations(range(q))
    ]


def fixed_point_table(
    words: Sequence[Sequence[int]], q: int, n: int
) -> list[dict[str, int | str]]:
    """Count fixed words through their uniquely induced symbol permutations."""
    fixed_counts: Counter[tuple[int, int, tuple[int, ...]]] = Counter()
    for word in words:
        for orientation in (1, -1):
            for offset in range(n):
                permutation = [-1] * q
                used_images = [False] * q
                consistent = True
                for index, target in enumerate(word):
                    source = word[(offset + orientation * index) % n]
                    image = permutation[source]
                    if image == -1:
                        if used_images[target]:
                            consistent = False
                            break
                        permutation[source] = target
                        used_images[target] = True
                    elif image != target:
                        consistent = False
                        break
                if consistent and -1 not in permutation:
                    fixed_counts[
                        (orientation, offset, tuple(permutation))
                    ] += 1

    rows: list[dict[str, int | str]] = []
    permutations = list(itertools.permutations(range(q)))
    for orientation in (1, -1):
        for offset in range(n):
            for permutation in permutations:
                rows.append(
                    {
                        "orientation": orientation,
                        "offset": offset,
                        "permutation": "".join(map(str, permutation)),
                        "fixed_valid_words": fixed_counts[
                            (orientation, offset, permutation)
                        ],
                    }
                )
    return rows


def solve(instance_path: Path, output_path: Path) -> None:
    instance = parse_instance(instance_path)
    q: int = instance["q"]
    n: int = instance["n"]
    normalized_words, representatives = enumerate_normalized_words(instance)

    labelings = math.factorial(q)
    group_order = 2 * n * labelings
    words = labeled_words(normalized_words, q)
    orbit_rows: list[dict[str, int | str]] = []
    histogram: Counter[str] = Counter()
    for representative in sorted(representatives):
        dihedral_orbit_size = len(normalized_dihedral_images(representative))
        orbit_size = labelings * dihedral_orbit_size
        stabilizer_size = group_order // orbit_size
        row = {
            "representative": canonical_string(representative),
            "orbit_size": orbit_size,
            "stabilizer_size": stabilizer_size,
        }
        orbit_rows.append(row)
        histogram[str(orbit_size)] += 1

    valid_labeled_words = len(words)
    if sum(row["orbit_size"] for row in orbit_rows) != valid_labeled_words:
        raise RuntimeError("orbit partition failed its size identity")
    if any(
        row["orbit_size"] * row["stabilizer_size"] != group_order
        for row in orbit_rows
    ):
        raise RuntimeError("orbit-stabilizer identity failed")
    fixed_points = fixed_point_table(words, q, n)
    burnside_fixed_sum = sum(row["fixed_valid_words"] for row in fixed_points)
    if burnside_fixed_sum != group_order * len(orbit_rows):
        raise RuntimeError("Burnside identity failed")

    output = {
        "alphabet_size": q,
        "length": n,
        "group_order": group_order,
        "valid_labeled_words": valid_labeled_words,
        "orbit_count": len(orbit_rows),
        "orbit_size_histogram": dict(
            sorted(histogram.items(), key=lambda item: int(item[0]))
        ),
        "orbits": orbit_rows,
        "burnside_fixed_sum": burnside_fixed_sum,
        "fixed_point_table": fixed_points,
    }
    output_path.parent.mkdir(parents=True, exist_ok=True)
    output_path.write_text(
        json.dumps(output, indent=2, ensure_ascii=False) + "\n", encoding="utf-8"
    )


def main() -> None:
    parser = argparse.ArgumentParser()
    parser.add_argument("instance", type=Path)
    parser.add_argument("output", type=Path)
    args = parser.parse_args()
    solve(args.instance, args.output)


if __name__ == "__main__":
    main()
Evaluation ContractCategory, tested runtime, effort, and required artifacts+
Public identifierevaluation/weave-orbit-census
CategoryMathematics and Formal Reasoning
SubcategoryCombinatorics and enumeration
Expert estimate8 hours
Model testedOpus-4.8
Agent testedTerminus-2
Required artifacts/app/weave_orbits.py, /app/output.json
Difficulty DesignWhy the benchmark discriminates between plausible and correct work+

The task combines exact cyclic constraint enumeration with two interacting quotient operations and a complete Burnside certificate. A correct solver must normalize arbitrary symbol permutations, preserve cyclic boundary contributions at four distances and three incommensurate motif strides, canonicalize reflection-sensitive equality motifs, compute full orbit and stabilizer sizes without double-counting, and count fixed valid words for every group element in a prescribed order. The five-symbol shipped instance makes assignment-by-assignment labeled enumeration especially unattractive: its balanced multiplicity partition creates many indistinguishable frequency assignments, its symbol action alone has 120 elements, and the complete group has 3,600 elements acting on 10,800 valid labeled words in multiple orbits. The intended insight is to combine restricted-growth enumeration, partial motif-ledger pruning, and uniquely induced symbol permutations under each positional symmetry. The instance is a synthetic, hand-constructed research benchmark representative of computer-assisted enumeration audits. Researchers in enumerative combinatorics and combinatorics on words would use this census to validate orbit-generation software, diagnose symmetry handling, and obtain reproducible certificates for downstream classification.

Reference ApproachThe intended solution strategy+

The reference solution enumerates proper restricted-growth words, which chooses one representative of every full-symbol-relabeling class. It prunes prefixes using the sorted multiplicity partition, the number of unseen symbols, non-wrapping distance-match upper bounds, and exact partial motif ledgers: each cyclic stride window is charged as soon as its final indexed position is assigned, so an overfull canonical motif class rejects the branch immediately. Normalized rotations and reflections yield each canonical representative and its dihedral orbit size. After expanding normalized words through all symbol permutations, the solution constructs the Burnside table without scanning every group element against every word. For each word and positional dihedral image, full alphabet usage induces at most one consistent symbol permutation, whose ledger entry is incremented. Orbit-partition, orbit-stabilizer, and Burnside identities are asserted before serialization.

Verification DesignHow the produced artifacts are independently checked+

The verifier independently reconstructs exact censuses and full fixed-point ledgers from the shipped instance and generated hidden instances. It executes the submitted program on the shipped five-symbol, four-distance, three-stride case and requires both the regenerated result and the submitted output.json to match the exact census. Hidden coverage combines fixed diagnostic regressions with a deterministic self-binding challenge seed: a fixed domain constant is hashed with the finalized /app tree, so identical submissions receive identical cases while adding a precomputed lookup changes the cases it would need to answer. The resulting family includes private five-symbol cases and varies length, multiplicity partition, distances, motif geometry, orbit counts, and stabilizer spectra. A direct labeled-word model cross-checks the independently pruned normalized census on every smaller case. Tests compare the complete JSON structure, all representatives, every group-element row, and every exact integer; validate prescribed ordering and all three group identities; execute submitted code with a scrubbed environment, reduced privileges, and a per-process timeout; reject link-based artifacts; and hash-pin protected input bytes. No numerical tolerances are used.

Task 16 · End-to-end software build

Recover Echo Step

Agentic software engineering · Model Training and ML Infrastructure · Training loops

Recover a stateful low-precision optimizer policy from archived training traces.

Full brief · 7 source files · 1,158 lines · 8 expert hours

Task BriefThe complete agent-facing instruction+

The implementation of an in-house low-precision optimizer policy was lost. Archived executions are available in /app/data/archives.json; each entry contains an input document and the exact result produced by the retired policy. The archive is the normative behavioral specification. /app/data/FORMAT.md defines the input and result schemas.

Reimplement the policy in /app/echo_train.py. It must support:

python3 /app/echo_train.py INPUT_JSON OUTPUT_JSON

For every valid input in the documented domain, replace OUTPUT_JSON with the corresponding result. Your implementation must derive the policy from the archives rather than special-case the shipped documents. Fresh verification documents use different parameter names, vector shapes, groups, numerical settings, event histories, accumulation windows, cadence phases, clipping regimes, and overflow positions.

Run the program on the input of archive a01 and leave that result at /app/output.json. Do not modify /app/data.

The completed artifacts must satisfy:

  1. /app/echo_train.py and /app/output.json are regular files, and the CLI replaces an existing output file.
  2. Every archived execution is reproduced within the numerical comparison described in FORMAT.md.
  3. Fresh executions follow the same accumulation, transaction, optimizer-state, cadence, clipping, quantization, and overflow policy evidenced by the archive.
  4. Results have exactly the documented schema, event ordering, and finite JSON numbers.
Implementation & Verification Code7 authored source files: solution, public tests, and build environment+
Source Browsersolution/generate_archives.py
7 files · 1,158 lines
import json
from pathlib import Path

from solve import run_document


def cfg(target=2, quantum=0.25, decay=1.0, groups=None):
    return {
        "target_tokens": target,
        "echo_decay": decay,
        "groups": groups or {"g": {"clip_norm": 100.0}},
        "quantum": quantum,
    }


def param(
    name,
    value,
    group="g",
    beta=0.0,
    lr=0.1,
    weight_decay=0.0,
    period=1,
    phase=1,
):
    return {
        "name": name,
        "group": group,
        "value": value,
        "beta": beta,
        "lr": lr,
        "weight_decay": weight_decay,
        "period": period,
        "phase": phase,
    }


def batch(tokens, gradients, overflow=False):
    return {
        "type": "batch",
        "tokens": tokens,
        "overflow": overflow,
        "grad_sums": gradients,
    }


def flush(identifier):
    return {"type": "flush", "id": identifier}


def doc(config, parameters, events):
    return {
        "config": config,
        "parameters": parameters,
        "events": events,
    }


def documents():
    return [
        doc(
            cfg(target=4),
            [param("weight", [1.0, -1.0])],
            [
                batch(2, {"weight": [1.0, -0.5]}),
                flush("wait"),
                batch(2, {"weight": [1.4, 1.5]}),
                flush("commit"),
            ],
        ),
        doc(
            cfg(target=1, quantum=0.1),
            [param("weight", [2.0], beta=0.75, lr=0.2, weight_decay=0.05)],
            [
                batch(1, {"weight": [0.6]}),
                flush("one"),
                batch(1, {"weight": [-0.2]}),
                flush("two"),
                batch(1, {"weight": [0.4]}),
                flush("three"),
            ],
        ),
        doc(
            cfg(target=1, quantum=0.25),
            [param("ties", [0.0, 0.0, 0.0, 0.0])],
            [
                batch(1, {"ties": [0.125, 0.375, -0.125, -0.375]}),
                flush("ties"),
                batch(1, {"ties": [0.125, -0.125, 0.125, -0.125]}),
                flush("carry"),
            ],
        ),
        doc(
            cfg(target=1, quantum=0.2, decay=0.5),
            [param("decayed", [0.5, -0.5], beta=0.2)],
            [
                batch(1, {"decayed": [0.31, -0.49]}),
                flush("one"),
                batch(1, {"decayed": [0.07, 0.11]}),
                flush("two"),
                batch(1, {"decayed": [-0.23, 0.29]}),
                flush("three"),
            ],
        ),
        doc(
            cfg(
                target=2,
                quantum=0.5,
                groups={"limited": {"clip_norm": 2.5}},
            ),
            [param("vector", [1.0, 1.0], group="limited")],
            [
                batch(2, {"vector": [5.2, 6.8]}),
                flush("clipped"),
                batch(2, {"vector": [0.4, -0.6]}),
                flush("echoed"),
            ],
        ),
        doc(
            cfg(
                target=1,
                quantum=0.25,
                groups={"shared": {"clip_norm": 1.1}},
            ),
            [
                param("left", [1.0, -1.0], group="shared", lr=0.05),
                param("right", [0.5], group="shared", lr=0.2),
            ],
            [
                batch(1, {"left": [1.2, -0.8], "right": [1.6]}),
                flush("joint"),
                batch(1, {"left": [0.3, 0.7], "right": [-0.4]}),
                flush("joint-echo"),
            ],
        ),
        doc(
            cfg(
                target=1,
                quantum=0.2,
                groups={
                    "a": {"clip_norm": 0.5},
                    "b": {"clip_norm": 3.0},
                },
            ),
            [
                param("alpha", [0.0], group="a"),
                param("beta", [0.0, 0.0], group="b"),
            ],
            [
                batch(1, {"alpha": [1.3], "beta": [1.0, -2.0]}),
                flush("separate"),
            ],
        ),
        doc(
            cfg(target=2, quantum=0.25),
            [
                param("fast", [0.5], period=1, phase=1),
                param("slow", [1.0], beta=0.5, period=2, phase=2),
            ],
            [
                batch(2, {"fast": [1.0], "slow": [2.0]}),
                flush("step-1"),
                batch(2, {"fast": [0.5], "slow": [1.0]}),
                flush("step-2"),
                batch(2, {"fast": [-0.5], "slow": [3.0]}),
                flush("step-3"),
            ],
        ),
        doc(
            cfg(target=2, quantum=0.5),
            [
                param("fast", [0.0], period=1, phase=1),
                param("slow", [0.0], period=3, phase=3),
            ],
            [
                batch(2, {"fast": [1.0], "slow": [2.0]}),
                flush("commit-1"),
                batch(2, {"fast": [50.0], "slow": [60.0]}, overflow=True),
                flush("overflow"),
                batch(2, {"fast": [3.0], "slow": [4.0]}),
                flush("commit-2"),
                batch(2, {"fast": [1.0], "slow": [6.0]}),
                flush("commit-3"),
            ],
        ),
        doc(
            cfg(target=4, quantum=0.25),
            [param("weight", [1.0])],
            [
                batch(1, {"weight": [0.5]}),
                batch(1, {"weight": [20.0]}, overflow=True),
                flush("waiting-overflow"),
                batch(2, {"weight": [1.5]}),
                flush("abort"),
                batch(4, {"weight": [2.0]}),
                flush("recovery"),
            ],
        ),
        doc(
            cfg(target=1, quantum=0.25),
            [
                param("present", [0.0], beta=0.5),
                param("silent", [1.0], beta=0.5),
            ],
            [
                batch(1, {"present": [0.6]}),
                flush("missing-is-zero"),
                batch(1, {"silent": [0.4]}),
                flush("roles-reverse"),
            ],
        ),
        doc(
            cfg(target=3, quantum=0.125, decay=0.75),
            [
                param("matrix", [1.0, -1.0, 0.5], beta=0.4, lr=0.03),
                param("scalar", [2.0], beta=0.8, lr=0.07),
            ],
            [
                batch(5, {"matrix": [2.3, -1.1, 0.7], "scalar": [3.2]}),
                flush("excess"),
                batch(1, {"matrix": [0.2, 0.3, -0.4]}),
            ],
        ),
        doc(
            cfg(target=1, quantum=0.2),
            [
                param("p1", [0.0], period=3, phase=1),
                param("p2", [0.0], period=3, phase=2),
                param("p3", [0.0], period=3, phase=3),
            ],
            [
                batch(1, {"p1": [0.4], "p2": [0.8], "p3": [1.2]}),
                flush("one"),
                batch(1, {"p1": [0.5], "p2": [0.7], "p3": [0.9]}),
                flush("two"),
                batch(1, {"p1": [0.6], "p2": [0.6], "p3": [0.6]}),
                flush("three"),
                batch(1, {"p1": [0.7], "p2": [0.5], "p3": [0.3]}),
                flush("four"),
            ],
        ),
        doc(
            cfg(
                target=2,
                quantum=0.25,
                decay=0.6,
                groups={"shared": {"clip_norm": 0.9}},
            ),
            [
                param(
                    "fast",
                    [0.5, -0.25],
                    group="shared",
                    beta=0.6,
                    lr=0.08,
                    weight_decay=0.02,
                ),
                param(
                    "slow",
                    [1.0],
                    group="shared",
                    beta=0.3,
                    lr=0.04,
                    weight_decay=0.01,
                    period=2,
                    phase=2,
                ),
            ],
            [
                batch(2, {"fast": [1.3, -0.7], "slow": [1.1]}),
                flush("one"),
                batch(1, {"fast": [4.0, 4.0]}, overflow=True),
                flush("waiting"),
                batch(1, {"slow": [5.0]}),
                flush("overflow"),
                batch(2, {"fast": [0.2, 1.4], "slow": [2.2]}),
                flush("two"),
                batch(3, {"fast": [-1.7, 0.4], "slow": [-0.8]}),
                flush("three"),
                batch(2, {"fast": [1.1, -1.5], "slow": [0.6]}),
                flush("four"),
            ],
        ),
    ]


def main():
    archive = []
    for index, document in enumerate(documents(), start=1):
        archive.append({
            "id": f"a{index:02d}",
            "input": document,
            "result": run_document(document),
        })
    destination = Path(__file__).parents[1] / "environment" / "data" / "archives.json"
    destination.write_text(
        json.dumps(archive, allow_nan=False, indent=2) + "\n",
        encoding="utf-8",
        newline="\n",
    )


if __name__ == "__main__":
    main()
Evaluation ContractCategory, tested runtime, effort, and required artifacts+
Public identifierevaluation/recover-echo-step
CategoryModel Training and ML Infrastructure
SubcategoryTraining loops
Expert estimate8 hours
Model testedOpus-4.8
Agent testedTerminus-2
Required artifacts/app/echo_train.py, /app/output.json
Difficulty DesignWhy the benchmark discriminates between plausible and correct work+

The task requires reconstructing a custom low-precision training-loop policy from behavioral evidence rather than recalling a named optimizer. The decisive interaction is clip-aware error feedback maintained in pre-clipping space: quantized transmitted gradients, group clipping, residual decay, multi-rate parameter escrow, velocity, decoupled weight decay, and overflow recovery form a noncommutative state transition. The archive isolates every primitive but fresh streams combine them, so familiar error-feedback, post-clip residual, whole-window reset, or globally-cadenced implementations remain plausible on simple traces and diverge only after dependent updates. This is realistic optimizer-kernel recovery work for senior training-infrastructure engineers and requires iterative numerical hypothesis testing, implementation, and differential validation.

Reference ApproachThe intended solution strategy+

The reference reconstructs each accumulation window, preserving per-parameter escrow across successful transactions until that parameter's cadence is due. For due parameters it combines the escrow mean with the persistent echo, computes due-only group clipping, quantizes the clipped signal, maps quantization error back through the clip factor before applying echo decay, then advances velocity, decoupled weight decay, values, and local clocks. Overflow clears only the current window; it does not erase older parameter escrow or optimizer state. The implementation emits post-event audit state for every flush.

Verification DesignHow the produced artifacts are independently checked+

The verifier hash-pins the agent-visible archive, then executes the submitted CLI in an unprivileged isolated directory against all archived pairs, fixed interaction cases, and submission-bound generated streams with different identities, shapes, groups, cadences, clipping factors, ties, and overflow locations. An independent functional reference model computes fresh expected results without importing the submitted or oracle implementation. Parsed JSON is compared recursively: schemas and discrete values exactly, finite floats with relative tolerance 1e-9 and absolute tolerance 1e-10. The case set was mutation-audited against post-clip residuals, overflow-cleared escrow, and global rather than parameter-local cadence; each shortcut diverges on normative or fresh evidence.

Task 17 · End-to-end software build

Cairnpack Archive

Agentic software engineering · File and Media Operations · Archiving and compression

Build a canonical sparse archive using globally optimized bounded-depth delta recipes

Full brief · 6 source files · 2,256 lines · 7 expert hours

Task BriefThe complete agent-facing instruction+

The preservation catalog in /app/data/catalog.json describes a filesystem image with sparse regular files, hard links, symbolic links, and directories. Build a reusable Python 3 archiver at /app/cairnpack.py that implements the complete CairnPack 1 contract in /app/data/FORMAT.md.

The program must support both commands below:

python3 /app/cairnpack.py pack INPUT_DIR ARCHIVE REPORT

python3 /app/cairnpack.py unpack ARCHIVE OUTPUT_DIR

Packing must validate the source catalog, perform the specified content-defined chunking and deduplication, choose the canonical globally optimal bounded-depth delta forest, and emit the binary archive and JSON report defined by the contract. Packing invalid input must exit nonzero without replacing an existing archive or report. Unpacking must validate the complete archive before replacing OUTPUT_DIR, then recreate its filesystem image, including sparse extents, link relationships, modes, and file contents. It must reject malformed archives without replacing an existing output directory.

Run the pack command on the supplied catalog:

python3 /app/cairnpack.py pack /app/data /app/output.cpk /app/report.json

Leave /app/cairnpack.py, /app/output.cpk, and /app/report.json in place as regular files, not symbolic links. The implementation will also be evaluated with fresh valid and invalid catalogs and archives governed by /app/data/FORMAT.md. Use only the Python standard library and do not modify /app/data.

Within the contract, a regular source file excludes symbolic links and any source path that traverses a symbolic-link component, even if resolution would remain beneath the command's input directory.

Implementation & Verification Code6 authored source files: solution, public tests, and build environment+
Source Browsersolution/cairnpack.py
6 files · 2,256 lines
#!/usr/bin/env python3
import binascii
import collections
import hashlib
import json
import os
import shutil
import stat
import struct
import sys
import tempfile
import unicodedata
from pathlib import Path


MAGIC = b"CAIRNPK1"
METHODS = ("RAW", "RLE", "XRLE", "XSPARSE")


class FormatError(ValueError):
    pass


def exact(obj, keys, label):
    if not isinstance(obj, dict) or set(obj) != set(keys):
        raise FormatError(f"{label} has invalid fields")


def integer(value, low, high, label):
    if isinstance(value, bool) or not isinstance(value, int) or not low <= value <= high:
        raise FormatError(f"{label} is out of range")
    return value


def canonical_json(value):
    return json.dumps(
        value, sort_keys=True, separators=(",", ":"), ensure_ascii=False
    ).encode("utf-8")


def uvarint(value):
    integer(value, 0, 2**63 - 1, "varint")
    out = bytearray()
    while True:
        byte = value & 0x7F
        value >>= 7
        if value:
            out.append(byte | 0x80)
        else:
            out.append(byte)
            return bytes(out)


def read_uvarint(data, pos):
    start = pos
    value = 0
    shift = 0
    for _ in range(10):
        if pos >= len(data):
            raise FormatError("truncated varint")
        byte = data[pos]
        pos += 1
        value |= (byte & 0x7F) << shift
        if not byte & 0x80:
            if value > 2**63 - 1 or data[start:pos] != uvarint(value):
                raise FormatError("noncanonical varint")
            return value, pos
        shift += 7
    raise FormatError("oversized varint")


def packbits(raw):
    out = bytearray()
    pos = 0
    while pos < len(raw):
        run = 1
        while pos + run < len(raw) and raw[pos + run] == raw[pos] and run < 130:
            run += 1
        if run >= 3:
            out.extend((0x80 | (run - 3), raw[pos]))
            pos += run
            continue
        end = pos
        while end < len(raw) and end - pos < 128:
            future = 1
            while (
                end + future < len(raw)
                and raw[end + future] == raw[end]
                and future < 3
            ):
                future += 1
            if future >= 3:
                break
            end += 1
        literal = raw[pos:end]
        if not literal:
            raise AssertionError("empty literal")
        out.append(len(literal) - 1)
        out.extend(literal)
        pos = end
    return bytes(out)


def unpackbits(payload):
    out = bytearray()
    pos = 0
    while pos < len(payload):
        control = payload[pos]
        pos += 1
        if control & 0x80:
            if pos >= len(payload):
                raise FormatError("truncated RLE run")
            out.extend(bytes((payload[pos],)) * ((control & 0x7F) + 3))
            pos += 1
        else:
            length = control + 1
            if pos + length > len(payload):
                raise FormatError("truncated RLE literal")
            out.extend(payload[pos : pos + length])
            pos += length
    if packbits(bytes(out)) != payload:
        raise FormatError("noncanonical RLE")
    return bytes(out)


def xsparse(base, raw):
    changes = [(i, a ^ b) for i, (a, b) in enumerate(zip(base, raw)) if a != b]
    out = bytearray(uvarint(len(changes)))
    previous = -1
    for position, value in changes:
        out.extend(uvarint(position - previous - 1))
        out.append(value)
        previous = position
    return bytes(out)


def apply_xsparse(base, payload):
    count, pos = read_uvarint(payload, 0)
    out = bytearray(base)
    previous = -1
    for _ in range(count):
        gap, pos = read_uvarint(payload, pos)
        position = previous + gap + 1
        if position >= len(out) or pos >= len(payload) or payload[pos] == 0:
            raise FormatError("invalid XSPARSE change")
        out[position] ^= payload[pos]
        pos += 1
        previous = position
    if pos != len(payload) or xsparse(base, bytes(out)) != payload:
        raise FormatError("noncanonical XSPARSE")
    return bytes(out)


def valid_path(path):
    if (
        not isinstance(path, str)
        or not path
        or "\x00" in path
        or path.startswith("/")
        or path.endswith("/")
        or unicodedata.normalize("NFC", path) != path
        or len(path.encode("utf-8")) > 160
    ):
        return False
    return all(part not in ("", ".", "..") for part in path.split("/"))


def safe_source(root, source):
    if not valid_path(source) or source == "catalog.json":
        raise FormatError("invalid source path")
    root_real = Path(root).resolve()
    candidate = root_real
    for component in source.split("/"):
        candidate /= component
        if candidate.is_symlink():
            raise FormatError("source path traverses a symlink")
    candidate = candidate.resolve()
    try:
        candidate.relative_to(root_real)
    except ValueError as exc:
        raise FormatError("source escapes input") from exc
    if not candidate.is_file():
        raise FormatError("source is not a regular file")
    return candidate


def validate_catalog(root):
    catalog_path = Path(root) / "catalog.json"
    try:
        raw = catalog_path.read_bytes()
        catalog = json.loads(raw.decode("utf-8"))
    except (OSError, UnicodeError, json.JSONDecodeError) as exc:
        raise FormatError("invalid catalog JSON") from exc
    exact(catalog, ("format", "chunker", "optimizer", "entries"), "catalog")
    if catalog["format"] != "cairn-source-1":
        raise FormatError("invalid source format")

    chunker = catalog["chunker"]
    exact(chunker, ("minimum", "mask", "target", "maximum"), "chunker")
    minimum = integer(chunker["minimum"], 4, 64, "minimum")
    mask = integer(chunker["mask"], 1, 255, "mask")
    target = integer(chunker["target"], 0, mask, "target")
    maximum = integer(chunker["maximum"], minimum, 256, "maximum")
    if mask & (mask + 1):
        raise FormatError("mask is not 2^k-1")

    optimizer = catalog["optimizer"]
    exact(
        optimizer,
        ("max_depth", "fanout", "delta_budget", "restart_interval"),
        "optimizer",
    )
    integer(optimizer["max_depth"], 1, 2, "max_depth")
    integer(optimizer["fanout"], 1, 4, "fanout")
    integer(optimizer["delta_budget"], 0, 64, "delta_budget")
    integer(optimizer["restart_interval"], 2, 32, "restart_interval")

    entries = catalog["entries"]
    if not isinstance(entries, list) or not 1 <= len(entries) <= 64:
        raise FormatError("invalid entry count")
    normalized = []
    paths = set()
    file_content = {}
    path_types = {}
    for source_entry in entries:
        if not isinstance(source_entry, dict):
            raise FormatError("entry is not an object")
        path = source_entry.get("path")
        if not valid_path(path) or path in paths:
            raise FormatError("invalid or duplicate archive path")
        paths.add(path)
        kind = source_entry.get("type")
        if kind == "directory":
            exact(source_entry, ("path", "type", "mode"), "directory")
            entry = dict(source_entry)
            integer(entry["mode"], 0, 4095, "mode")
        elif kind == "symlink":
            exact(source_entry, ("path", "type", "target"), "symlink")
            target_value = source_entry["target"]
            if not isinstance(target_value, str) or not target_value or "\x00" in target_value:
                raise FormatError("invalid symlink target")
            entry = dict(source_entry)
        elif kind == "hardlink":
            exact(source_entry, ("path", "type", "target"), "hardlink")
            if not isinstance(source_entry["target"], str):
                raise FormatError("invalid hardlink target")
            entry = dict(source_entry)
        elif kind == "file":
            exact(
                source_entry,
                ("path", "type", "mode", "inode", "size", "extents"),
                "file",
            )
            mode = integer(source_entry["mode"], 0, 4095, "mode")
            size = integer(source_entry["size"], 0, 1048576, "size")
            inode = source_entry["inode"]
            if (
                not isinstance(inode, str)
                or not inode
                or unicodedata.normalize("NFC", inode) != inode
                or len(inode.encode("utf-8")) > 80
            ):
                raise FormatError("invalid inode")
            extents = source_entry["extents"]
            if not isinstance(extents, list):
                raise FormatError("extents is not an array")
            previous_end = 0
            loaded = []
            semantic = []
            for index, extent in enumerate(extents):
                exact(extent, ("offset", "length", "source"), "extent")
                offset = integer(extent["offset"], 0, size, "extent offset")
                length = integer(extent["length"], 1, size or 1, "extent length")
                if (index and offset < previous_end) or offset + length > size:
                    raise FormatError("overlapping or out-of-range extent")
                source_path = safe_source(root, extent["source"])
                data = source_path.read_bytes()
                if len(data) != length:
                    raise FormatError("extent source length mismatch")
                previous_end = offset + length
                loaded.append({"offset": offset, "length": length, "data": data})
                semantic.append((offset, length, data))
            signature = (mode, size, tuple(semantic))
            if inode in file_content and file_content[inode] != signature:
                raise FormatError("inconsistent repeated inode")
            file_content[inode] = signature
            entry = {
                "path": path,
                "type": "file",
                "mode": mode,
                "inode": inode,
                "size": size,
                "_extents": loaded,
            }
        else:
            raise FormatError("invalid entry type")
        normalized.append(entry)
        path_types[path] = kind

    normalized.sort(key=lambda item: item["path"].encode("utf-8"))
    order = {entry["path"]: index for index, entry in enumerate(normalized)}
    for entry in normalized:
        path = entry["path"]
        if "/" in path:
            parent = path.rsplit("/", 1)[0]
            if path_types.get(parent) != "directory":
                raise FormatError("missing directory parent")
        if entry["type"] == "hardlink":
            target = entry["target"]
            if path_types.get(target) != "file" or order[target] >= order[path]:
                raise FormatError("invalid hardlink target")
    return normalized, chunker, optimizer


def foldgear(data, config):
    chunks = []
    start = 0
    h = 0
    for pos, byte in enumerate(data):
        h = ((h << 5) ^ (h >> 2) ^ ((byte * 40503 + 277) & 65535)) & 65535
        length = pos - start + 1
        if (
            length >= config["minimum"] and (h & config["mask"]) == config["target"]
        ) or length == config["maximum"]:
            chunks.append(data[start : pos + 1])
            start = pos + 1
            h = 0
    if start < len(data):
        chunks.append(data[start:])
    return chunks


def build_manifest(entries, chunker):
    chunks = []
    digest_to_ordinal = {}
    manifest_entries = []
    for entry in entries:
        kind = entry["type"]
        if kind != "file":
            manifest_entries.append(dict(entry))
            continue
        out_extents = []
        for extent in entry["_extents"]:
            ordinals = []
            for raw in foldgear(extent["data"], chunker):
                digest = hashlib.sha256(raw).digest()
                if digest in digest_to_ordinal:
                    ordinal = digest_to_ordinal[digest]
                    if chunks[ordinal] != raw:
                        raise FormatError("digest collision")
                else:
                    ordinal = len(chunks)
                    digest_to_ordinal[digest] = ordinal
                    chunks.append(raw)
                ordinals.append(ordinal)
            out_extents.append(
                {
                    "offset": extent["offset"],
                    "length": extent["length"],
                    "chunks": ordinals,
                }
            )
        manifest_entries.append(
            {
                "path": entry["path"],
                "type": "file",
                "mode": entry["mode"],
                "inode": entry["inode"],
                "size": entry["size"],
                "extents": out_extents,
            }
        )
    manifest = {"format": "cairnpack-manifest-1", "entries": manifest_entries}
    length_counts = collections.Counter(map(len, chunks))
    if len(chunks) > 64 or any(count > 12 for count in length_counts.values()):
        raise FormatError("chunk population exceeds catalog bounds")
    return manifest, chunks


def optimize(chunks, config):
    roots = []
    candidates = []
    for ordinal, raw in enumerate(chunks):
        encoded = packbits(raw)
        if len(raw) <= len(encoded):
            root = (0, -1, raw, len(raw) + 1)
        else:
            root = (1, -1, encoded, len(encoded) + 1)
        roots.append(root)
        options = [root]
        if ordinal % config["restart_interval"]:
            for base in range(ordinal):
                if len(chunks[base]) != len(raw):
                    continue
                xor = bytes(a ^ b for a, b in zip(chunks[base], raw))
                for method, payload in ((2, packbits(xor)), (3, xsparse(chunks[base], raw))):
                    cost = len(payload) + len(uvarint(base + 1))
                    if cost < root[3]:
                        options.append((method, base, payload, cost))
        options.sort(key=lambda item: (item[0], item[1] + 1))
        candidates.append(options)

    suffix = [0] * (len(chunks) + 1)
    for index in range(len(chunks) - 1, -1, -1):
        suffix[index] = suffix[index + 1] + min(item[3] for item in candidates[index])

    best_cost = sum(item[3] for item in roots)
    best_signature = tuple((item[0], 0) for item in roots)
    best = list(roots)
    selected = []
    depths = []
    child_count = [0] * len(chunks)

    def search(index, cost, delta_count):
        nonlocal best_cost, best_signature, best
        if cost + suffix[index] > best_cost:
            return
        if index == len(chunks):
            signature = tuple((method, base + 1) for method, base, _, _ in selected)
            if cost < best_cost or (cost == best_cost and signature < best_signature):
                best_cost = cost
                best_signature = signature
                best = list(selected)
            return
        for option in candidates[index]:
            method, base, payload, option_cost = option
            if method >= 2:
                if delta_count >= config["delta_budget"]:
                    continue
                depth = depths[base] + 1
                if depth > config["max_depth"] or child_count[base] >= config["fanout"]:
                    continue
                child_count[base] += 1
            else:
                depth = 0
            selected.append(option)
            depths.append(depth)
            search(index + 1, cost + option_cost, delta_count + (method >= 2))
            depths.pop()
            selected.pop()
            if method >= 2:
                child_count[base] -= 1

    search(0, 0, 0)
    result = []
    depth_values = []
    for method, base, payload, cost in best:
        depth = 0 if base < 0 else depth_values[base] + 1
        depth_values.append(depth)
        result.append(
            {
                "method": method,
                "base": base,
                "payload": payload,
                "depth": depth,
                "cost": cost,
            }
        )
    return result


def build_archive(manifest, chunks, recipes):
    archive = bytearray(MAGIC)
    manifest_bytes = canonical_json(manifest)
    archive.extend(uvarint(len(manifest_bytes)))
    archive.extend(manifest_bytes)
    archive.extend(uvarint(len(chunks)))
    index = []
    for ordinal, (raw, recipe) in enumerate(zip(chunks, recipes)):
        start = len(archive)
        digest = hashlib.sha256(raw).digest()
        archive.extend(digest)
        archive.extend(uvarint(len(raw)))
        archive.append(recipe["method"])
        archive.extend(uvarint(recipe["base"] + 1))
        archive.extend(uvarint(len(recipe["payload"])))
        archive.extend(recipe["payload"])
        archive.extend(struct.pack("<I", binascii.crc32(raw) & 0xFFFFFFFF))
        index.append(
            {
                "digest": digest.hex(),
                "length": len(archive) - start,
                "offset": start,
            }
        )
    index_bytes = canonical_json(index)
    archive.extend(b"IDX1")
    archive.extend(uvarint(len(index_bytes)))
    archive.extend(index_bytes)
    archive.extend(b"END1")
    archive.extend(hashlib.sha256(archive).digest())
    return bytes(archive)


def report_for(entries, chunks, recipes, archive):
    depth_max = max((recipe["depth"] for recipe in recipes), default=0)
    report = {
        "format": "cairnpack-report-1",
        "entry_count": len(entries),
        "file_count": sum(entry["type"] == "file" for entry in entries),
        "logical_bytes": sum(
            entry.get("size", 0) for entry in entries if entry["type"] == "file"
        ),
        "stored_chunk_count": len(chunks),
        "unique_bytes": sum(map(len, chunks)),
        "encoded_payload_bytes": sum(len(recipe["payload"]) for recipe in recipes),
        "delta_chunks": sum(recipe["method"] >= 2 for recipe in recipes),
        "max_delta_depth": depth_max,
        "archive_sha256": hashlib.sha256(archive).hexdigest(),
        "recipes": [],
    }
    for ordinal, (raw, recipe) in enumerate(zip(chunks, recipes)):
        report["recipes"].append(
            {
                "ordinal": ordinal,
                "digest": hashlib.sha256(raw).hexdigest(),
                "method": METHODS[recipe["method"]],
                "base_ordinal": None if recipe["base"] < 0 else recipe["base"],
                "depth": recipe["depth"],
                "raw_size": len(raw),
                "payload_size": len(recipe["payload"]),
            }
        )
    return report


def atomic_pack(archive_path, report_path, archive, report):
    archive_path = Path(archive_path)
    report_path = Path(report_path)
    if archive_path == report_path:
        raise FormatError("pack destinations must differ")
    archive_path.parent.mkdir(parents=True, exist_ok=True)
    report_path.parent.mkdir(parents=True, exist_ok=True)
    archive_tmp = tempfile.NamedTemporaryFile(
        dir=archive_path.parent, prefix=".cairn-", delete=False
    )
    report_tmp = tempfile.NamedTemporaryFile(
        dir=report_path.parent, prefix=".cairn-", delete=False
    )
    try:
        archive_tmp.write(archive)
        archive_tmp.flush()
        os.fsync(archive_tmp.fileno())
        archive_tmp.close()
        report_tmp.write(canonical_json(report) + b"\n")
        report_tmp.flush()
        os.fsync(report_tmp.fileno())
        report_tmp.close()
        os.replace(archive_tmp.name, archive_path)
        os.replace(report_tmp.name, report_path)
    finally:
        archive_tmp.close()
        report_tmp.close()
        for path in (archive_tmp.name, report_tmp.name):
            try:
                os.unlink(path)
            except FileNotFoundError:
                pass


def parse_manifest(value, chunk_count):
    exact(value, ("format", "entries"), "manifest")
    if value["format"] != "cairnpack-manifest-1" or not isinstance(value["entries"], list):
        raise FormatError("invalid manifest")
    paths = {}
    inodes = {}
    previous_path = None
    for entry in value["entries"]:
        if not isinstance(entry, dict):
            raise FormatError("invalid manifest entry")
        kind = entry.get("type")
        path = entry.get("path")
        if not valid_path(path) or (
            previous_path is not None
            and path.encode("utf-8") <= previous_path.encode("utf-8")
        ):
            raise FormatError("noncanonical manifest path order")
        previous_path = path
        if "/" in path and paths.get(path.rsplit("/", 1)[0]) != "directory":
            raise FormatError("missing manifest parent")
        if kind == "directory":
            exact(entry, ("path", "type", "mode"), "manifest directory")
            integer(entry["mode"], 0, 4095, "mode")
        elif kind == "symlink":
            exact(entry, ("path", "type", "target"), "manifest symlink")
            if (
                not isinstance(entry["target"], str)
                or not entry["target"]
                or "\x00" in entry["target"]
            ):
                raise FormatError("invalid manifest symlink")
        elif kind == "hardlink":
            exact(entry, ("path", "type", "target"), "manifest hardlink")
            if paths.get(entry["target"]) != "file":
                raise FormatError("invalid manifest hardlink")
        elif kind == "file":
            exact(
                entry,
                ("path", "type", "mode", "inode", "size", "extents"),
                "manifest file",
            )
            integer(entry["mode"], 0, 4095, "mode")
            size = integer(entry["size"], 0, 1048576, "size")
            inode = entry["inode"]
            if (
                not isinstance(inode, str)
                or not inode
                or unicodedata.normalize("NFC", inode) != inode
                or len(inode.encode("utf-8")) > 80
                or not isinstance(entry["extents"], list)
            ):
                raise FormatError("invalid manifest inode or extents")
            prior_end = 0
            signature = []
            for index, extent in enumerate(entry["extents"]):
                exact(extent, ("offset", "length", "chunks"), "manifest extent")
                offset = integer(extent["offset"], 0, size, "offset")
                length = integer(extent["length"], 1, size or 1, "length")
                chunks = extent["chunks"]
                if (
                    (index and offset < prior_end)
                    or offset + length > size
                    or not isinstance(chunks, list)
                    or any(
                        isinstance(item, bool)
                        or not isinstance(item, int)
                        or not 0 <= item < chunk_count
                        for item in chunks
                    )
                ):
                    raise FormatError("invalid manifest extent")
                prior_end = offset + length
                signature.append((offset, length, tuple(chunks)))
            inode_signature = (entry["mode"], size, tuple(signature))
            if inode in inodes and inodes[inode] != inode_signature:
                raise FormatError("inconsistent manifest inode")
            inodes[inode] = inode_signature
        else:
            raise FormatError("invalid manifest entry type")
        paths[path] = kind
    return value["entries"]


def parse_archive(data):
    if len(data) < 8 + 1 + 1 + 4 + 1 + 4 + 32 or not data.startswith(MAGIC):
        raise FormatError("invalid archive header")
    if hashlib.sha256(data[:-32]).digest() != data[-32:]:
        raise FormatError("archive trailer digest mismatch")
    if data[-36:-32] != b"END1":
        raise FormatError("missing archive trailer")
    limit = len(data) - 36
    pos = len(MAGIC)
    manifest_length, pos = read_uvarint(data, pos)
    if pos + manifest_length > limit:
        raise FormatError("truncated manifest")
    manifest_bytes = data[pos : pos + manifest_length]
    pos += manifest_length
    try:
        manifest = json.loads(manifest_bytes.decode("utf-8"))
    except (UnicodeError, json.JSONDecodeError) as exc:
        raise FormatError("invalid manifest JSON") from exc
    if canonical_json(manifest) != manifest_bytes:
        raise FormatError("noncanonical manifest JSON")
    chunk_count, pos = read_uvarint(data, pos)
    if chunk_count > 100000:
        raise FormatError("excessive chunk count")
    chunks = []
    recipes = []
    index = []
    depths = []
    digests_seen = set()
    for ordinal in range(chunk_count):
        start = pos
        if pos + 32 > limit:
            raise FormatError("truncated record digest")
        digest = data[pos : pos + 32]
        pos += 32
        raw_length, pos = read_uvarint(data, pos)
        if raw_length > 256 or pos >= limit:
            raise FormatError("invalid chunk length")
        method = data[pos]
        pos += 1
        if method not in range(4):
            raise FormatError("invalid method")
        base_value, pos = read_uvarint(data, pos)
        payload_length, pos = read_uvarint(data, pos)
        if pos + payload_length + 4 > limit:
            raise FormatError("truncated record")
        payload = data[pos : pos + payload_length]
        pos += payload_length
        crc = struct.unpack("<I", data[pos : pos + 4])[0]
        pos += 4
        base = base_value - 1
        if method < 2:
            if base_value != 0:
                raise FormatError("root recipe has a base")
            raw = payload if method == 0 else unpackbits(payload)
            depth = 0
            root_encoded = packbits(raw)
            canonical_method = 0 if len(raw) <= len(root_encoded) else 1
            if method != canonical_method:
                raise FormatError("noncanonical root recipe")
        else:
            if not 0 <= base < ordinal:
                raise FormatError("invalid delta base")
            base_raw = chunks[base]
            if method == 2:
                xor = unpackbits(payload)
                if len(xor) != len(base_raw):
                    raise FormatError("XRLE length mismatch")
                raw = bytes(a ^ b for a, b in zip(base_raw, xor))
            else:
                raw = apply_xsparse(base_raw, payload)
            depth = depths[base] + 1
            if depth > 2:
                raise FormatError("delta depth exceeds format limit")
        if (
            len(raw) != raw_length
            or (binascii.crc32(raw) & 0xFFFFFFFF) != crc
            or hashlib.sha256(raw).digest() != digest
            or digest in digests_seen
        ):
            raise FormatError("record integrity mismatch")
        digests_seen.add(digest)
        chunks.append(raw)
        depths.append(depth)
        recipes.append(
            {"method": method, "base": base, "payload": payload, "depth": depth}
        )
        index.append(
            {"digest": digest.hex(), "length": pos - start, "offset": start}
        )
    if data[pos : pos + 4] != b"IDX1":
        raise FormatError("missing index")
    pos += 4
    index_length, pos = read_uvarint(data, pos)
    if pos + index_length != limit:
        raise FormatError("invalid index boundary")
    index_bytes = data[pos : pos + index_length]
    try:
        supplied_index = json.loads(index_bytes.decode("utf-8"))
    except (UnicodeError, json.JSONDecodeError) as exc:
        raise FormatError("invalid index JSON") from exc
    if canonical_json(supplied_index) != index_bytes or supplied_index != index:
        raise FormatError("noncanonical or incorrect index")
    entries = parse_manifest(manifest, chunk_count)
    first_seen = set()
    for entry in entries:
        if entry["type"] == "file":
            for extent in entry["extents"]:
                if sum(len(chunks[item]) for item in extent["chunks"]) != extent["length"]:
                    raise FormatError("extent chunk lengths do not match")
                for ordinal in extent["chunks"]:
                    if ordinal not in first_seen:
                        if ordinal != len(first_seen):
                            raise FormatError("noncanonical chunk first-occurrence order")
                        first_seen.add(ordinal)
    if len(first_seen) != chunk_count:
        raise FormatError("unreferenced chunk record")
    return manifest, entries, chunks, recipes


def replace_tree(temp_path, destination):
    destination = Path(destination)
    destination.parent.mkdir(parents=True, exist_ok=True)
    backup = None
    if destination.exists() or destination.is_symlink():
        backup = destination.parent / (
            f".{destination.name}.old-{next(tempfile._get_candidate_names())}"
        )
        os.replace(destination, backup)
    try:
        os.replace(temp_path, destination)
    except Exception:
        if backup is not None:
            os.replace(backup, destination)
        raise
    if backup is not None:
        if backup.is_dir() and not backup.is_symlink():
            shutil.rmtree(backup)
        else:
            backup.unlink()


def materialize(entries, chunks, destination):
    destination = Path(destination)
    destination.parent.mkdir(parents=True, exist_ok=True)
    temp = Path(tempfile.mkdtemp(dir=destination.parent, prefix=".cairn-tree-"))
    inode_paths = {}
    entry_paths = {}
    try:
        for entry in entries:
            if entry["type"] == "directory":
                path = temp.joinpath(*entry["path"].split("/"))
                path.mkdir()
                entry_paths[entry["path"]] = path
        for entry in entries:
            kind = entry["type"]
            path = temp.joinpath(*entry["path"].split("/"))
            if kind == "file":
                if entry["inode"] in inode_paths:
                    os.link(inode_paths[entry["inode"]], path)
                else:
                    with open(path, "wb") as handle:
                        handle.truncate(entry["size"])
                        for extent in entry["extents"]:
                            handle.seek(extent["offset"])
                            for ordinal in extent["chunks"]:
                                handle.write(chunks[ordinal])
                    os.chmod(path, entry["mode"])
                    inode_paths[entry["inode"]] = path
                entry_paths[entry["path"]] = path
            elif kind == "hardlink":
                os.link(entry_paths[entry["target"]], path)
                entry_paths[entry["path"]] = path
            elif kind == "symlink":
                os.symlink(entry["target"], path)
                entry_paths[entry["path"]] = path
        for entry in reversed(entries):
            if entry["type"] == "directory":
                os.chmod(entry_paths[entry["path"]], entry["mode"])
        replace_tree(temp, destination)
    except Exception:
        if temp.exists():
            shutil.rmtree(temp)
        raise


def command_pack(input_dir, archive_path, report_path):
    entries, chunker, optimizer = validate_catalog(input_dir)
    manifest, chunks = build_manifest(entries, chunker)
    recipes = optimize(chunks, optimizer)
    archive = build_archive(manifest, chunks, recipes)
    report = report_for(entries, chunks, recipes, archive)
    atomic_pack(archive_path, report_path, archive, report)


def command_unpack(archive_path, output_dir):
    data = Path(archive_path).read_bytes()
    _, entries, chunks, _ = parse_archive(data)
    materialize(entries, chunks, output_dir)


def main(argv):
    try:
        if len(argv) == 5 and argv[1] == "pack":
            command_pack(argv[2], argv[3], argv[4])
        elif len(argv) == 4 and argv[1] == "unpack":
            command_unpack(argv[2], argv[3])
        else:
            raise FormatError("invalid command line")
    except (FormatError, OSError, UnicodeError, json.JSONDecodeError) as exc:
        print(f"cairnpack: {exc}", file=sys.stderr)
        return 2
    return 0


if __name__ == "__main__":
    raise SystemExit(main(sys.argv))
Evaluation ContractCategory, tested runtime, effort, and required artifacts+
Public identifierevaluation/cairnpack-archive
CategoryFile and Media Operations
SubcategoryArchiving and compression
Expert estimate7 hours
Model testedOpus-4.8
Agent testedTerminus-2
Required artifacts/app/cairnpack.py, /app/output.cpk, /app/report.json
Difficulty DesignWhy the benchmark discriminates between plausible and correct work+

The task combines preservation-grade filesystem semantics, a byte-exact seekable container, canonical content-defined chunking, two custom delta encodings, and a coupled global recipe optimization. Fan-out, depth, restart, and budget constraints make greedy compression incorrect; an expert must reason about a constrained dependency forest while treating exhaustive schema, path, link, extent, archive, and transactional validation as a co-equal correctness problem. The catalog and CairnPack format are synthetic, deterministic fixtures designed to expose interactions rather than imitate a published codec. The intended practitioner is a backup-format or digital-preservation engineer building a reproducible archival container that must retain sparse layout and Unix link identity while balancing compression ratio against bounded random-access dependencies.

Reference ApproachThe intended solution strategy+

The reference validates the exact source schema, safe paths, link order, repeated-inode consistency, extent bounds, and source lengths before canonicalizing entries. It chunks every sparse extent with the FoldGear state machine, deduplicates by digest in first-occurrence order, and derives RAW, RLE, XRLE, and XSPARSE candidates from the supplied bytes. A branch-and-bound search minimizes actual recipe wire cost under the stated restart, depth, fan-out, and delta-budget constraints, using the complete recipe signature for ties. It serializes the canonical manifest, chunk records, absolute seek index, integrity trailer, and report through temporary sibling files. Unpack first validates canonical encodings, offsets, references, checksums, digests, and safe manifest paths, then materializes sparse files and link relationships in a complete sibling tree before replacing the destination.

Verification DesignHow the produced artifacts are independently checked+

The verifier exercises the submitted CLI on the supplied catalog and on a fixed-seed family of sixteen procedurally generated held-out catalogs outside /app/data. Their object bytes, path namespaces, modes, inode labels, sparse geometry, symlink targets, chunk sizes, and optimizer parameters vary across the family, making a narrow implementation for the two original fixtures insufficient while keeping clean runs exactly reproducible; depth-one and depth-two configurations are both covered. A separate paired source proves that nondefault FoldGear mask and target values materially alter chunk boundaries. Each held-out source is independently modeled and byte-snapshotted before untrusted code runs, and the snapshot must remain unchanged, so agent-writable inputs cannot redefine expected results. Before supplied inputs contribute to expected results, the verifier checks the exact file set and SHA-256 of every catalog, specification, and object. An independent parser reconstructs every record and recomputes FoldGear boundaries, deduplication order, all four recipe payloads, XSPARSE selection, the globally optimal forest and tie-break, canonical manifest/index/report bytes, checksums, and archive digests. Unpack checks require the exact destination path set, logical contents, physical sparse holes, modes, symlink targets, shared identity within each repeated-inode group, and distinct identities across different groups. A bounded large-file case observes an existing destination during unpack and rejects publication before the complete replacement tree is visible. Positive replacement checks cover existing pack destinations, and a chunkless catalog pins the empty-recipe report edge. Parameterized negative tests use the same generated family and cover missing and extra fields across every schema level, exact ranges, path normalization and traversal, forbidden catalog self-sourcing, final-component and intermediate-component source symlinks, escaping sources, entry and chunk population limits, inode and extent invariants, invalid link target types, noncanonical varints, trailing archive bytes, damaged indexes and records, unsafe extraction, unsupported commands, and transactional preservation of every present or absent destination. The submitted CLI runs under isolated Python without site packages and with a per-process timeout. JSON artifacts are compared with exact scalar types as well as values, so booleans cannot impersonate integer fields. All comparisons are exact because the contract fully specifies integer arithmetic, byte encodings, ordering, and JSON serialization.

Task 18 · End-to-end software build

Nacre NPT Replay

Agentic software engineering · Systems Infrastructure and Operations · Virtualization and emulation

Reconstruct nested virtual-memory behavior across stale translation caches, DMA, and snapshots.

Full brief · 6 source files · 2,263 lines · 14 expert hours

Task BriefThe complete agent-facing instruction+

Implement the Nacre-NPT nested-memory replay monitor defined by the normative contract in /app/data/PROTOCOL.md.

Create a self-contained Python program at /app/nacre_npt.py with this interface:

python3 /app/nacre_npt.py INPUT_JSON OUTPUT_JSON

The program must accept every valid Nacre-NPT document described by the contract and atomically replace OUTPUT_JSON with the complete replay result. The monitor couples bounded set-associative guest, nested, and IOMMU caches with non-coherent map edits, prefix-acknowledged shootdowns, atomic compare-and-swap, globally ordered scoped DMA barriers, host dirty logging, vCPU rebinding, permission faults, and branchable guest snapshots. All transition policies, fault precedence, pending-work boundaries, canonical arrays, and JSON schemas in /app/data/PROTOCOL.md are normative.

Run the program on /app/data/sample.json and leave its result at /app/output.json. Fresh valid documents vary cache geometry, page size, mappings, topology, eviction history, dirty-tracker overlap, pending work, permissions, and snapshot branches.

/app/nacre_npt.py and /app/output.json must be regular files with one filesystem link each. Derive the replay entirely from INPUT_JSON, use only the Python standard library, and leave /app/data unchanged.

Implementation & Verification Code6 authored source files: solution, public tests, and build environment+
Source Browsersolution/nacre_npt.py
6 files · 2,263 lines
#!/usr/bin/env python3
"""Reference implementation of the Nacre-NPT 1 replay contract."""

from __future__ import annotations

import copy
import json
import os
import sys
import tempfile
from pathlib import Path
from typing import Any


class AccessFault(Exception):
    """A contract-defined virtual-memory access fault."""

    def __init__(self, code: str):
        super().__init__(code)
        self.code = code


Entry = dict[str, Any]
Table = dict[tuple[str, int], Entry]


def expand_table(
    spans: list[dict[str, Any]],
    scope_name: str,
    source_name: str,
    target_name: str,
) -> Table:
    """Expand non-overlapping input spans into authoritative base-page entries."""
    table: Table = {}
    for span in spans:
        for delta in range(span["pages"]):
            table[(span[scope_name], span[source_name] + delta)] = {
                "target": span[target_name] + delta,
                "perms": span["perms"],
            }
    return table


class Monitor:
    """Mutable machine state and deterministic replay operations."""

    def __init__(self, document: dict[str, Any]):
        self.page_size = document["page_size"]
        self.geometry = document["cache_geometry"]
        self.clock = 0
        self.frames = {
            frame["hpa_page"]: bytearray.fromhex(frame["data"])
            for frame in document["frames"]
        }
        self.vcpus = {
            row["id"]: {"asid": row["asid"], "vmid": row["vmid"]}
            for row in document["vcpus"]
        }
        self.devices = {
            row["id"]: {"domain": row["domain"]} for row in document["devices"]
        }
        self.guest = expand_table(
            document["guest_mappings"], "asid", "va_page", "gpa_page"
        )
        self.nested = expand_table(
            document["nested_mappings"], "vmid", "gpa_page", "hpa_page"
        )
        self.iommu = expand_table(
            document["iommu_mappings"], "domain", "iova_page", "hpa_page"
        )
        self.guest_tlb: dict[tuple[str, str, int], Entry] = {}
        self.nested_tlb: dict[tuple[str, str, int], Entry] = {}
        self.iotlb: dict[tuple[str, str, int], Entry] = {}
        self.pending: list[dict[str, Any]] = []
        self.shootdowns = {vcpu: [] for vcpu in self.vcpus}
        self.dirty_trackers: dict[str, dict[str, Any]] = {}
        self.snapshots: dict[str, dict[str, Any]] = {}
        self.transcript: list[dict[str, Any]] = []

    @staticmethod
    def require(entry: Entry, access: str, stage: str) -> None:
        if access not in entry["perms"]:
            raise AccessFault(f"{stage}_{'read' if access == 'r' else 'write'}_denied")

    def cache_walk(
        self,
        table: Table,
        cache: dict[tuple[str, str, int], Entry],
        tier: str,
        owner: str,
        scope: str,
        page: int,
        access: str,
    ) -> int:
        """Walk one bounded cache, updating LRU before permission evaluation."""
        cache_key = (owner, scope, page)
        entry = cache.get(cache_key)
        if entry is None:
            authoritative = table.get((scope, page))
            if authoritative is None:
                raise AccessFault(f"{tier}_not_present")
            sets = self.geometry[tier]["sets"]
            ways = self.geometry[tier]["ways"]
            set_index = page % sets
            occupants = [
                key
                for key in cache
                if key[0] == owner and key[2] % sets == set_index
            ]
            if len(occupants) == ways:
                victim = min(occupants, key=lambda key: cache[key]["last_used"])
                del cache[victim]
            entry = authoritative.copy()
            cache[cache_key] = entry
        self.clock += 1
        entry["last_used"] = self.clock
        self.require(entry, access, tier)
        return entry["target"]

    def guest_walk(self, vcpu: str, va_page: int, access: str) -> int:
        binding = self.vcpus[vcpu]
        return self.cache_walk(
            self.guest,
            self.guest_tlb,
            "guest",
            vcpu,
            binding["asid"],
            va_page,
            access,
        )

    def nested_walk(self, vcpu: str, gpa_page: int, access: str) -> int:
        binding = self.vcpus[vcpu]
        return self.cache_walk(
            self.nested,
            self.nested_tlb,
            "nested",
            vcpu,
            binding["vmid"],
            gpa_page,
            access,
        )

    def cpu_translate(self, vcpu: str, address: int, access: str) -> tuple[int, int]:
        va_page, offset = divmod(address, self.page_size)
        gpa_page = self.guest_walk(vcpu, va_page, access)
        hpa_page = self.nested_walk(vcpu, gpa_page, access)
        return hpa_page, offset

    def iommu_walk(self, device: str, iova_page: int, access: str) -> int:
        return self.cache_walk(
            self.iommu,
            self.iotlb,
            "iommu",
            device,
            self.devices[device]["domain"],
            iova_page,
            access,
        )

    def dma_translate(self, device: str, address: int, access: str) -> tuple[int, int]:
        iova_page, offset = divmod(address, self.page_size)
        hpa_page = self.iommu_walk(device, iova_page, access)
        return hpa_page, offset

    @staticmethod
    def map_range(
        table: Table,
        scope: str,
        source: int,
        target: int,
        pages: int,
        perms: str,
    ) -> None:
        for delta in range(pages):
            table[(scope, source + delta)] = {
                "target": target + delta,
                "perms": perms,
            }

    @staticmethod
    def unmap_range(table: Table, scope: str, source: int, pages: int) -> None:
        for delta in range(pages):
            table.pop((scope, source + delta), None)

    @staticmethod
    def protect_range(
        table: Table, scope: str, source: int, pages: int, perms: str
    ) -> None:
        for delta in range(pages):
            table[(scope, source + delta)]["perms"] = perms

    def take_snapshot(self, name: str) -> None:
        self.snapshots[name] = copy.deepcopy(
            {
                "frames": self.frames,
                "vcpus": self.vcpus,
                "guest": self.guest,
                "nested": self.nested,
                "iommu": self.iommu,
                "pending": self.pending,
                "shootdowns": self.shootdowns,
            }
        )

    def restore_snapshot(self, name: str) -> None:
        restored = copy.deepcopy(self.snapshots[name])
        self.frames = restored["frames"]
        self.vcpus = restored["vcpus"]
        self.guest = restored["guest"]
        self.nested = restored["nested"]
        self.iommu = restored["iommu"]
        self.pending = restored["pending"]
        self.shootdowns = restored["shootdowns"]
        self.guest_tlb.clear()
        self.nested_tlb.clear()
        self.iotlb.clear()

    def mark_dirty(self, hpa_page: int) -> None:
        """Mark a committed host-page write in every overlapping live tracker."""
        for tracker in self.dirty_trackers.values():
            first = tracker["hpa_page"]
            if first <= hpa_page < first + tracker["pages"]:
                tracker["dirty"].add(hpa_page)

    @staticmethod
    def harvested(tracker: dict[str, Any]) -> dict[str, Any]:
        return {"pages": sorted(tracker["dirty"])}

    def apply(self, operation: dict[str, Any]) -> dict[str, Any]:
        op = operation["op"]

        if op == "cpu.read":
            hpa, offset = self.cpu_translate(operation["vcpu"], operation["address"], "r")
            return {"value": self.frames[hpa][offset]}
        if op == "cpu.write":
            hpa, offset = self.cpu_translate(operation["vcpu"], operation["address"], "w")
            self.frames[hpa][offset] = operation["value"]
            self.mark_dirty(hpa)
            return {}
        if op == "cpu.cas":
            hpa, offset = self.cpu_translate(operation["vcpu"], operation["address"], "w")
            observed = self.frames[hpa][offset]
            swapped = observed == operation["expected"]
            if swapped:
                self.frames[hpa][offset] = operation["value"]
                self.mark_dirty(hpa)
            return {"observed": observed, "swapped": swapped}
        if op == "dma.read":
            hpa, offset = self.dma_translate(
                operation["device"], operation["address"], "r"
            )
            return {"value": self.frames[hpa][offset]}
        if op == "dma.write":
            device = operation["device"]
            hpa, offset = self.dma_translate(device, operation["address"], "w")
            self.pending.append(
                {
                    "device": device,
                    "hpa_page": hpa,
                    "offset": offset,
                    "value": operation["value"],
                }
            )
            queued = sum(write["device"] == device for write in self.pending)
            return {"queued": queued}
        if op == "dma.fence":
            if "device" in operation:
                eligible = lambda write: write["device"] == operation["device"]
            elif "domain" in operation:
                eligible = lambda write: (
                    self.devices[write["device"]]["domain"] == operation["domain"]
                )
            else:
                eligible = lambda write: True
            remaining = []
            committed = 0
            for write in self.pending:
                if eligible(write):
                    self.frames[write["hpa_page"]][write["offset"]] = write["value"]
                    self.mark_dirty(write["hpa_page"])
                    committed += 1
                else:
                    remaining.append(write)
            self.pending = remaining
            return {"committed": committed}

        if op == "guest.map":
            self.map_range(
                self.guest,
                operation["asid"],
                operation["va_page"],
                operation["gpa_page"],
                operation["pages"],
                operation["perms"],
            )
            return {}
        if op == "guest.unmap":
            self.unmap_range(
                self.guest,
                operation["asid"],
                operation["va_page"],
                operation["pages"],
            )
            return {}
        if op == "guest.protect":
            self.protect_range(
                self.guest,
                operation["asid"],
                operation["va_page"],
                operation["pages"],
                operation["perms"],
            )
            return {}

        if op == "nested.map":
            self.map_range(
                self.nested,
                operation["vmid"],
                operation["gpa_page"],
                operation["hpa_page"],
                operation["pages"],
                operation["perms"],
            )
            return {}
        if op == "nested.unmap":
            self.unmap_range(
                self.nested,
                operation["vmid"],
                operation["gpa_page"],
                operation["pages"],
            )
            return {}
        if op == "nested.protect":
            self.protect_range(
                self.nested,
                operation["vmid"],
                operation["gpa_page"],
                operation["pages"],
                operation["perms"],
            )
            return {}

        if op == "iommu.map":
            self.map_range(
                self.iommu,
                operation["domain"],
                operation["iova_page"],
                operation["hpa_page"],
                operation["pages"],
                operation["perms"],
            )
            return {}
        if op == "iommu.unmap":
            self.unmap_range(
                self.iommu,
                operation["domain"],
                operation["iova_page"],
                operation["pages"],
            )
            return {}
        if op == "iommu.protect":
            self.protect_range(
                self.iommu,
                operation["domain"],
                operation["iova_page"],
                operation["pages"],
                operation["perms"],
            )
            return {}

        if op == "tlb.invlpg":
            vcpu = operation["vcpu"]
            asid = self.vcpus[vcpu]["asid"]
            va_page = operation["address"] // self.page_size
            self.guest_tlb.pop((vcpu, asid, va_page), None)
            return {}
        if op == "tlb.invguest":
            asid = operation["asid"]
            self.guest_tlb = {
                key: value for key, value in self.guest_tlb.items() if key[1] != asid
            }
            return {}
        if op == "tlb.invept":
            vmid = operation["vmid"]
            if "gpa_page" in operation:
                start = operation["gpa_page"]
                stop = start + operation["pages"]
                self.nested_tlb = {
                    key: value
                    for key, value in self.nested_tlb.items()
                    if not (key[1] == vmid and start <= key[2] < stop)
                }
            else:
                self.nested_tlb = {
                    key: value
                    for key, value in self.nested_tlb.items()
                    if key[1] != vmid
                }
            return {}
        if op == "tlb.post":
            vcpu = operation["vcpu"]
            self.shootdowns[vcpu].append(
                {
                    "ticket": operation["ticket"],
                    "vmid": operation["vmid"],
                    "gpa_page": operation["gpa_page"],
                    "pages": operation["pages"],
                }
            )
            return {"pending": len(self.shootdowns[vcpu])}
        if op == "tlb.ack":
            vcpu = operation["vcpu"]
            receipts = self.shootdowns[vcpu]
            if "through" in operation:
                applied = next(
                    index + 1
                    for index, receipt in enumerate(receipts)
                    if receipt["ticket"] == operation["through"]
                )
            else:
                applied = len(receipts)
            for receipt in receipts[:applied]:
                start = receipt["gpa_page"]
                stop = start + receipt["pages"]
                vmid = receipt["vmid"]
                self.nested_tlb = {
                    key: value
                    for key, value in self.nested_tlb.items()
                    if not (
                        key[0] == vcpu
                        and key[1] == vmid
                        and start <= key[2] < stop
                    )
                }
            self.shootdowns[vcpu] = receipts[applied:]
            return {"applied": applied}
        if op == "iotlb.invalidate":
            device = operation["device"]
            if "iova_page" in operation:
                page = operation["iova_page"]
                self.iotlb = {
                    key: value
                    for key, value in self.iotlb.items()
                    if not (key[0] == device and key[2] == page)
                }
            else:
                self.iotlb = {
                    key: value
                    for key, value in self.iotlb.items()
                    if key[0] != device
                }
            return {}
        if op == "vcpu.bind":
            self.vcpus[operation["vcpu"]] = {
                "asid": operation["asid"],
                "vmid": operation["vmid"],
            }
            return {}
        if op == "dirty.begin":
            self.dirty_trackers[operation["tracker"]] = {
                "hpa_page": operation["hpa_page"],
                "pages": operation["pages"],
                "dirty": set(),
            }
            return {}
        if op == "dirty.harvest":
            tracker = self.dirty_trackers[operation["tracker"]]
            result = self.harvested(tracker)
            tracker["dirty"].clear()
            return result
        if op == "dirty.end":
            tracker = self.dirty_trackers.pop(operation["tracker"])
            return self.harvested(tracker)
        if op == "snapshot.take":
            self.take_snapshot(operation["snapshot"])
            return {}
        if op == "snapshot.restore":
            self.restore_snapshot(operation["snapshot"])
            return {}
        if op == "snapshot.drop":
            del self.snapshots[operation["snapshot"]]
            return {}
        raise ValueError(f"unsupported operation: {op}")

    def replay(self, operations: list[dict[str, Any]]) -> None:
        for operation in operations:
            try:
                result = self.apply(operation)
                self.transcript.append(
                    {
                        "id": operation["id"],
                        "op": operation["op"],
                        "status": "ok",
                        "result": result,
                    }
                )
            except AccessFault as fault:
                self.transcript.append(
                    {
                        "id": operation["id"],
                        "op": operation["op"],
                        "status": "fault",
                        "fault": fault.code,
                    }
                )

    def final_state(self) -> dict[str, Any]:
        guest_mappings = [
            {
                "asid": asid,
                "va_page": page,
                "gpa_page": entry["target"],
                "perms": entry["perms"],
            }
            for (asid, page), entry in sorted(self.guest.items())
        ]
        nested_mappings = [
            {
                "vmid": vmid,
                "gpa_page": page,
                "hpa_page": entry["target"],
                "perms": entry["perms"],
            }
            for (vmid, page), entry in sorted(self.nested.items())
        ]
        iommu_mappings = [
            {
                "domain": domain,
                "iova_page": page,
                "hpa_page": entry["target"],
                "perms": entry["perms"],
            }
            for (domain, page), entry in sorted(self.iommu.items())
        ]
        guest_tlb = [
            {
                "vcpu": vcpu,
                "asid": asid,
                "va_page": page,
                "gpa_page": entry["target"],
                "perms": entry["perms"],
            }
            for (vcpu, asid, page), entry in sorted(self.guest_tlb.items())
        ]
        nested_tlb = [
            {
                "vcpu": vcpu,
                "vmid": vmid,
                "gpa_page": page,
                "hpa_page": entry["target"],
                "perms": entry["perms"],
            }
            for (vcpu, vmid, page), entry in sorted(self.nested_tlb.items())
        ]
        iotlb = [
            {
                "device": device,
                "domain": domain,
                "iova_page": page,
                "hpa_page": entry["target"],
                "perms": entry["perms"],
            }
            for (device, domain, page), entry in sorted(self.iotlb.items())
        ]
        pending_dma = [write.copy() for write in self.pending]
        pending_shootdowns = []
        for vcpu in sorted(self.shootdowns):
            for receipt in self.shootdowns[vcpu]:
                pending_shootdowns.append({"vcpu": vcpu, **receipt})
        dirty_trackers = [
            {
                "tracker": name,
                "hpa_page": tracker["hpa_page"],
                "pages": tracker["pages"],
                "dirty_pages": sorted(tracker["dirty"]),
            }
            for name, tracker in sorted(self.dirty_trackers.items())
        ]
        return {
            "frames": [
                {"hpa_page": page, "data": bytes(data).hex()}
                for page, data in sorted(self.frames.items())
            ],
            "vcpus": [
                {"id": vcpu, **self.vcpus[vcpu]} for vcpu in sorted(self.vcpus)
            ],
            "devices": [
                {"id": device, **self.devices[device]}
                for device in sorted(self.devices)
            ],
            "guest_mappings": guest_mappings,
            "nested_mappings": nested_mappings,
            "iommu_mappings": iommu_mappings,
            "guest_tlb": guest_tlb,
            "nested_tlb": nested_tlb,
            "iotlb": iotlb,
            "pending_dma": pending_dma,
            "pending_shootdowns": pending_shootdowns,
            "dirty_trackers": dirty_trackers,
            "snapshots": sorted(self.snapshots),
        }

    def result(self) -> dict[str, Any]:
        return {
            "format": "nacre-npt-1",
            "transcript": self.transcript,
            "final": self.final_state(),
        }


def solve(document: dict[str, Any]) -> dict[str, Any]:
    """Replay a valid Nacre-NPT input document."""
    monitor = Monitor(document)
    monitor.replay(document["operations"])
    return monitor.result()


def atomic_write_json(path: Path, result: dict[str, Any]) -> None:
    """Write a JSON document through an adjacent temporary file."""
    path.parent.mkdir(parents=True, exist_ok=True)
    descriptor, temporary = tempfile.mkstemp(prefix=f".{path.name}.", dir=path.parent)
    try:
        with os.fdopen(descriptor, "w", encoding="utf-8", newline="\n") as stream:
            json.dump(result, stream, ensure_ascii=False, separators=(",", ":"))
            stream.write("\n")
        os.replace(temporary, path)
    except BaseException:
        try:
            os.unlink(temporary)
        except FileNotFoundError:
            pass
        raise


def main(argv: list[str]) -> int:
    if len(argv) != 3:
        print(f"usage: {argv[0]} INPUT_JSON OUTPUT_JSON", file=sys.stderr)
        return 2
    input_path = Path(argv[1])
    output_path = Path(argv[2])
    with input_path.open("r", encoding="utf-8") as stream:
        document = json.load(stream)
    atomic_write_json(output_path, solve(document))
    return 0


if __name__ == "__main__":
    raise SystemExit(main(sys.argv))
Evaluation ContractCategory, tested runtime, effort, and required artifacts+
Public identifierevaluation/nacre-npt-replay
CategorySystems Infrastructure and Operations
SubcategoryVirtualization and emulation
Expert estimate14 hours
Model testedOpus-4.8
Agent testedTerminus-2
Required artifacts/app/nacre_npt.py, /app/output.json
Difficulty DesignWhy the benchmark discriminates between plausible and correct work+

The solver must coordinate three differently tagged, bounded set-associative translation caches with LRU replacement, non-coherent map edits, partial fills on faults, and context reuse. Deferred nested-TLB receipts capture old lineage, allow prefix acknowledgements after rebinding, and survive guest snapshot branches. Posted DMA is globally ordered but drained by device, domain, or global barriers; its commit order drives overlapping memory writes and live-migration dirty trackers that remain outside guest snapshots. The synthetic traces reproduce failures normally diagnosed by senior hypervisor, VMM, IOMMU, live-migration, and CPU-emulator engineers.

Reference ApproachThe intended solution strategy+

The reference expands authoritative maps per page and applies a common bounded-LRU cache walk for guest, nested, and IOMMU translations. It records DMA in one global issue queue, drains selected entries without reordering survivors, applies shootdown prefixes, implements compare-and-swap, and marks every overlapping host dirty tracker only when a write commits. Guest snapshots copy authoritative state and pending work while excluding caches, trackers, and the registry. Canonical serialization emits the transcript and complete live state.

Verification DesignHow the produced artifacts are independently checked+

The verifier compares parsed JSON with an independently structured replay on the sample and hidden histories. Tests combine cache pressure and denied fills, stale permissions, context reactivation, prefix shootdowns, compare-and-swap, overlapping dirty trackers, cross-device DMA barriers, and snapshot restoration. Exact full-state comparisons cover eviction, global queue ordering, pending work, tracker state, and lineage; filesystem and integrity checks protect the requested artifacts and inputs.

Task 19 · Multi-source research

Broadcast Archive Source Chain

Factual verification · primary-source research · adversarial retrieval

A multi-hop factual research task joining film awards, biography, broadcast history, and a primary-source yearbook.

Complete task package · 5 authored records

Submission PackagePrompt, answer, source chain, and compliance checklist+

the research evaluation program — Task 1 Submission Package

Prompt

At the 61st Academy Awards, an actress received a rare double nomination — Best Actress for a biographical drama and Best Supporting Actress for a romantic comedy. Her father was an executive at a major radio network during the 1940s. According to the 1947 Broadcasting Yearbook, in the table titled "National Networks' Gross Monthly Time Sales, 1927–1946," what were that network's gross monthly time sales for the month of January 1934?

Word count: 72


Answer

$2,891,667


Golden Trajectory

Step 1: Identify the actress

Search: actress nominated for both Best Actress and Best Supporting Actress same year Open: Wikipedia or Oscars.org page listing double acting nominees Find: Sigourney Weaver — nominated at the 61st Academy Awards (1989 ceremony) for Best Actress (Gorillas in the Mist) and Best Supporting Actress (Working Girl)

Step 2: Identify her father and his network

Search: Sigourney Weaver father Open: Wikipedia page for Sigourney Weaver or Pat Weaver Find: Her father is Sylvester L. "Pat" Weaver, an NBC radio/television executive during the 1940s-1950s. The network is NBC.

Step 3: Locate the 1947 Broadcasting Yearbook

Search: 1947 Broadcasting Yearbook NBC gross monthly time sales January 1934 Open: World Radio History page for the 1947 Broadcasting Yearbook Navigate to: The table titled "National Networks' Gross Monthly Time Sales, 1927–1946" Find: NBC column, January 1934 row → $2,891,667

Step 4: Confirm the answer

Answer: $2,891,667


Verification Sources

  1. Academy Awards double nomination — Sigourney Weaver's double nomination at the 61st Oscars https://en.wikipedia.org/wiki/List_of_actors_nominated_for_multiple_Academy_Awards_in_the_same_year — "Nominees" table, 1988 (61st) row: Sigourney Weaver, Best Actress (Gorillas in the Mist) and Best Supporting Actress (Working Girl)

  2. Pat Weaver — NBC executive — Father of Sigourney Weaver, NBC executive https://www.museum.tv/tv-encyclopedia-18/weaver-sylvester-pat — Museum of Broadcast Communications entry confirming Pat Weaver's NBC executive roles (1949–56) and listing Susan (Sigourney) Weaver as his child

  3. 1947 Broadcasting Yearbook — "National Networks' Gross Monthly Time Sales, 1927–1946" table https://www.worldradiohistory.com/Archive-BC-YB/1947/1947-BC-YB.pdf — Page showing NBC January 1934: $2,891,667


Test Results

| Run | Model Answer | Correct? | Failure Mode | |-----|-------------|----------|--------------| | 1 | $1,309,662 | ❌ | NBC Red-only from wrong source | | 2 | $1,309,662 | ❌ | NBC Red-only, insists it's the Yearbook figure | | 3 | $2,391,667 | ❌ | Substituted 1935 Broadcasting table for 1947 Yearbook |

Score: 0/3 correct — 3/3 FAIL


Compliance Checklist

  • [x] Factual question with one defensible answer: $2,891,667
  • [x] Atomic answer — single dollar figure
  • [x] No named entities that give away the answer
  • [x] Requires 3+ distinct sources from 3+ domains (Oscars/Wikipedia, Wikipedia, worldradiohistory.com)
  • [x] No yes/no, true/false
  • [x] No paywalls — all sources freely accessible
  • [x] Timeless — the 1947 Yearbook figure will never change
  • [x] 70-150 words (72 words)
  • [x] No external references, stands on its own
  • [x] Every factual claim verified
  • [x] TV Shows & Movies domain — chains from film actress to broadcasting history
Golden TrajectoryThe complete research path and source-level evidence+

Golden Trajectory — Task 1

Group 1: Identify the actress

  1. Search: actress nominated for both Best Actress and Best Supporting Actress same year — the top result is the Wikipedia page "List of actors nominated for multiple Academy Awards in the same year." Click that result.

  2. Fetch: https://en.wikipedia.org/wiki/List_of_actors_nominated_for_multiple_Academy_Awards_in_the_same_year — scroll to the "Nominees" table, locate the row labeled "1988 (61st)" under the Year column. That row shows Sigourney Weaver with Best Actress nomination for Gorillas in the Mist and Best Supporting Actress nomination for Working Girl. The body text above the table also states: "Five did not receive an Academy Award in either category: Sigourney Weaver (nominations for Gorillas in the Mist and Working Girl)."

  3. Verify: Sigourney Weaver received both a Best Actress and Best Supporting Actress nomination in the same year, matching the prompt's constraint.

  4. This helps filter the prompt because only a dozen actors in Oscar history have received double acting nominations in the same year, and Sigourney Weaver is the one whose father was a radio network executive.


Group 2: Identify her father and his network

  1. Search: Sigourney Weaver father radio network executive — the search results include the Museum of Broadcast Communications encyclopedia entry for Pat Weaver. Click the museum.tv result.

  2. Fetch: https://www.museum.tv/tv-encyclopedia-18/weaver-sylvester-pat — in the biographical header at the top of the page, the entry lists his roles as "vice president, vice chair, president, then chair, NBC. 1949-56" and under children lists "Trajan Victor Charles and Susan (Sigourney)," confirming Sigourney Weaver is his daughter.

  3. Verify: Pat Weaver was an NBC executive, confirming the network is NBC.

  4. This helps filter the prompt because Sigourney Weaver's father was Pat Weaver, and his network was NBC — identifying the specific network whose gross monthly time sales must be looked up in the 1947 Broadcasting Yearbook.


Group 3: Find the figure in the 1947 Broadcasting Yearbook

  1. Search: "1947 Broadcasting Yearbook" "National Networks' Gross Monthly Time Sales" — the top result is the PDF hosted on worldradiohistory.com. Click that result.

  2. Fetch: https://www.worldradiohistory.com/Archive-BC-YB/1947/1947-BC-YB.pdf — navigate to the table titled "National Networks' Gross Monthly Time Sales, 1927–1946." Locate the NBC column header, then scan down to the January 1934 row. The value at the intersection of the NBC column and the January 1934 row is $2,891,667.

  3. Verify: The 1947 Broadcasting Yearbook's "National Networks' Gross Monthly Time Sales, 1927–1946" table shows NBC's gross monthly time sales for January 1934 as $2,891,667, matching the prompt's exact table title, network, and month.

  4. This helps filter the prompt because this is the final answer — the prompt specifies the exact yearbook (1947), exact table, exact network (NBC), and exact month (January 1934), so no other figure can be correct.

Answer: $2,891,667

Research RecordSource discovery and validation notes+

Task 1 — Research Dump

Source 1: Wikipedia - "Men Against Fire" (Black Mirror episode)

URL: https://en.wikipedia.org/wiki/Men_Against_Fire

  • Black Mirror S3E5, aired October 21, 2016 on Netflix
  • Written by Charlie Brooker, directed by Jakob Verbruggen
  • Starring Malachi Kirby as Stripe Koinange
  • Title comes from S.L.A. Marshall's 1947 book "Men Against Fire: The Problem of Battle Command"
  • Marshall claimed over 70% of WWII soldiers did not fire their rifles
  • Brooker also read Dave Grossman's "On Killing" based on Marshall's work
  • Original 2010 concept was titled "Inbound" — featured Norwegian invasion disguised as alien attack
  • Prosthetics designer: Kristyan Mallett (nominated for 2017 Make-Up Artists & Hair Stylists Guild Award)
  • Filmed at a disused army barracks near London
  • Episode received 59% on Rotten Tomatoes from 22 reviews
  • The psychiatrist Arquette (Michael Kelly) directly quotes Marshall's statistics in dialogue
  • Real-world tech comparison: ARC4 augmented reality military headset by Applied Research Associates

Source 2: Wikipedia - S.L.A. Marshall

URL: https://en.wikipedia.org/wiki/S.L.A._Marshall

  • Full name: Samuel Lyman Atwood Marshall (July 18, 1900 – December 17, 1977)
  • Nickname: "Slam"
  • Born: Catskill, New York
  • Died: El Paso, Texas, aged 77
  • Buried: Fort Bliss National Cemetery, Section A, Grave 124
  • WWI: Enlisted November 28, 1917, 315th Engineer Battalion, 90th Infantry Division
  • Served in France from June 1918, promoted to sergeant
  • Fought at the Battle of Saint-Mihiel and the Meuse-Argonne Offensive
  • His company (A Company, 315th Engineers) lost 8 dead and 15 wounded out of 165 men
  • Commissioned in early 1919, remained in France assisting with demobilization
  • Interwar: Worked as newspaper reporter/editor for the El Paso Herald and The Detroit News
  • Covered the Spanish Civil War as a journalist
  • Published "Blitzkrieg: Armies on Wheels" in 1940 analyzing Wehrmacht tactics
  • WWII: Joined Army's Center of Military History, pioneered after-action review technique
  • First combat assignment: Battle of Makin (November 1943)
  • Post-WWII: Led project employing 200+ former German officers including Heinz Guderian and Franz Halder
  • Korean War: Recalled for 3 months in late 1950 as Historian/Operations Analyst for Eighth Army
  • Vietnam War: Spent late 1966–early 1967 teaching after-action interview techniques
  • Co-authored "The Vietnam Primer" with Colonel David Hackworth
  • Retired from Army Reserve in 1960 with rank of Brigadier General
  • Wrote 30+ books including: "Men Against Fire" (1947), "Pork Chop Hill" (1956), "Bastogne" (1946), "The River and the Gauntlet" (1951), "Night Drop" (1962), "The Soldier's Load and The Mobility of a Nation" (1950)
  • Co-founded the Association of the U.S. Army
  • His after-action review technique remains standard military practice today
  • Grandson John Douglas Marshall was disowned after registering as conscientious objector during Vietnam War, later wrote memoir "Reconciliation Road" (1993)
  • Medals: Legion of Merit, Bronze Star (with Oak Leaf Cluster), French Croix de Guerre with Palm, Combat Infantryman Badge, among others

The Ratio of Fire Controversy (from same Wikipedia page)

  • Marshall claimed fewer than 25% of individual riflemen fired at an exposed enemy during WWII
  • Critics: David Hackworth called him "a voyeur rather than a warrior" and "a liar and profiteer"
  • Harold Leinbaugh (WWII infantry veteran): Called conclusions "absurd, ridiculous and totally nonsensical"
  • Roger Spiller (1988): Claimed Marshall's interviews either did not exist or were fabricated
  • Robert Engen (Canadian military historian): Marshall "wilfully disregarded important evidence"
  • Defenders: Dave Grossman argues the fundamental conclusion is confirmed by data from other armies and eras
  • John Keegan: Marshall's purpose was to "persuade the American army it was fighting its wars the wrong way"
  • FBI studies in 1950s–60s confirmed non-firing rates among law enforcement officers
  • Omer Bartov argued Wehrmacht's better combat performance partly resulted from deliberate brutalisation

Source 3: Roger Spiller's 1988 Article — "S.L.A. Marshall and the Ratio of Fire"

URL: https://doi.org/10.1080/03071848808445332 PDF: https://gwern.net/doc/history/s-l-a-marshall/1988-spiller.pdf

  • Published in RUSI Journal, Winter 1988, Vol. 133, No. 4, pp. 63–71
  • Spiller was Deputy Director of the Combat Studies Institute, U.S. Army Command and General Staff College
  • Key findings:
    • Marshall's surviving field notebooks show NO statistical compilations supporting the ratio
    • John Westover (Marshall's assistant in Europe) never recalled Marshall asking soldiers about firing weapons
    • Logistically impossible: 400 company interviews would have taken until late 1946 to complete
    • Marshall "exaggerated his own combat record" — never commanded infantry despite implying otherwise
    • Marshall treated statistics as "an adornment of belief" rather than a tool of inquiry
  • Spiller's balanced conclusion: Marshall was "one of the most important commentators on the soldier's world in this century" but the ratio of fire was "an invention"
  • Subsequent corroboration: John Whiteclay Chambers (2003) published new evidence from Lt. Frank Brennan confirming Marshall took only "cryptic notes"
  • Robert Engen (2011) in Canadian Military History found British/Canadian weapons studies uncovered no evidence of widespread non-firing

Source 4: Charlie Brooker Interview — Entertainment Weekly

URL: https://ew.com/article/2016/10/23/black-mirror-postmortem-interview-season-3/

  • Brooker read "On Killing" by Dave Grossman — called it "a bit of cheerful holiday reading"
  • "The stuff Michael Kelly is saying, how few soldiers want to pull the trigger, is true"
  • "I was reading about how people who dropped firebombs on Dresden didn't particularly suffer psychological consequences even though they knew they were burning people to death. Whereas if you have to slide a bayonet into somebody's ribs that stays with you forever"
  • Original 2010 draft was titled "Inbound" — rejected at the time as "a bit heavy-handed"
  • The very first draft involved an apparent alien attack on Britain later revealed as a Norwegian invasion
  • Inspired by John Pilger's documentary "The War You Don't See" (2010) about the Iraq War
  • Also inspired by reading a news article that referred to Middle Eastern immigrants as "cockroaches"

Source 5: Vox — "Men Against Fire" Review & Real-World Connections

URL: https://www.vox.com/2016/10/21/13327162/black-mirror-episode-5-men-against-fire-recap-review

  • Episode compared by multiple critics to Nazi Germany and the Holocaust
  • References to propaganda depicting enemies as insects
  • Real-world AR military headset: ARC4 by Applied Research Associates
  • Waverly Labs translation earpiece also referenced as real-world parallel
  • The episode explores how technology can dehumanize targets, making killing easier
  • Marshall's ratio of fire claim: "75% of soldiers did not fire at an exposed enemy"
  • The psychiatrist Arquette directly quotes Marshall's statistics in his dialogue with Stripe

Source 6: Collider — Iraq War Connection

URL: https://collider.com/black-mirror-episode-men-against-fire-iraq-war/

  • The episode has "eerie connections" to the Iraq War
  • Brooker was influenced by the dehumanizing language used in real conflicts
  • The "roach" terminology mirrors real-world propaganda techniques
  • The episode's themes connect to how soldiers are psychologically conditioned

Potential Prompt Chains

Chain A: TV → Military History → Biography → Geography

  1. Identify Black Mirror episode from content clues
  2. Find the 1947 book the title comes from
  3. Find the author (S.L.A. Marshall)
  4. Find where Marshall worked as a journalist (El Paso Herald)
  5. Find a specific fact about El Paso → FINAL ANSWER

Chain B: TV → Military History → Controversy → Academic Journal

  1. Identify episode from clues
  2. Find the book and author
  3. Find that Marshall's claims were debunked by Roger Spiller in 1988
  4. Find the journal Spiller published in (RUSI Journal)
  5. Find a specific fact about that journal → FINAL ANSWER

Chain C: TV → Military History → WWI → Battle Detail

  1. Identify episode from clues
  2. Find the book and author
  3. Find Marshall's WWI service — 315th Engineer Battalion, 90th Infantry Division
  4. Find the battle where his company lost 8 dead and 15 wounded (Saint-Mihiel or Meuse-Argonne)
  5. Find a specific fact about that battle → FINAL ANSWER
Reviewer ContextWhy the question is valid, atomic, and discriminating+

Additional Context for Reviewer — Task 1

1) Prompt is valid, self-contained, and unambiguous

Every constraint is pinned to a specific verifiable fact: double acting nomination in the same year (only 12 actors in Oscar history, fewer actresses), father as a 1940s radio network executive (biographical filter), and the exact 1947 Broadcasting Yearbook table title and month. No vague language, no time-sensitive anchors, no external references. Resolves to exactly one actress, one network, one figure.

2) Requires at least 3 URLs across 3 domains

  • wikipedia.org — List of actors with multiple same-year nominations → identifies Sigourney Weaver
  • museum.tv — Pat Weaver entry → confirms NBC executive role and father-daughter relationship
  • worldradiohistory.com — 1947 Broadcasting Yearbook PDF → the table cell: NBC × January 1934 = $2,891,667

Each hop depends on the prior one; the yearbook is useless without first identifying the network.

3) Stump is valid

All 3 runs failed with different wrong answers, and each failure is the model's, not the prompt's:

  • Run 1: $1,309,662 — read NBC Red-only column instead of combined NBC. Model even calculated Red+Blue ($2,298,165) but all three numbers were wrong.
  • Run 2: $1,309,662 — same Red-only misread, cross-referenced a contemporaneous table as false confirmation.
  • Run 3: $2,391,667 — substituted data from Broadcasting's 1935 table instead of the 1947 Yearbook.

The trap combines a scope trap (Red vs combined NBC column) with an adjacent-substitution trap (different yearbook editions contain different figures). The prompt's specificity means there's only one defensible answer; the model's failure to extract it is a genuine reading error.

4) Single defensible, unambiguous answer

$2,891,667 is atomic and verifiable in one PDF cell. No other figure can satisfy all constraints: the prompt specifies the 1947 Yearbook (not 1935), the exact table title, the exact network (NBC), and the exact month (January 1934). No rounding, no alternates, no interpretation ambiguity.

5) Additional notes

Cross-domain chain spans film/awards → biography → broadcasting history → industry trade publication. All sources are freely accessible, stable, and timeless — no logins, paywalls, or changing content.

Test ResultsObserved model attempts and documented failure modes+

Task 1 — Test Results

Prompt

An actress nominated for both Best Actress and Best Supporting Actress in the same year had a father who was an executive at a major radio network during the 1940s. According to the 1947 Broadcasting Yearbook, what were that network's gross monthly time sales for January 1934?

Correct Answer: $2,891,667


Run 1 — FAIL ❌

Model's Answer: $1,309,662

Model's Response:

The actress was Sigourney Weaver, who was nominated at the 61st Academy Awards for both Best Actress (Gorillas in the Mist) and Best Supporting Actress (Working Girl). Her father, Sylvester L. "Pat" Weaver, was an NBC executive, so the network is NBC.

According to the 1947 Broadcasting Yearbook's NBC network lineage — NBC-WEAF/Red — the gross monthly time sales for January 1934 were $1,309,662.

A small caveat: contemporaneous tables also list NBC-WJZ/Blue at $988,503 for January 1934; combining Red + Blue would give $2,298,165, but the NBC/Red network figure is $1,309,662.

Failure Mode: Model used NBC Red-only figure from a different source instead of the combined NBC figure in the 1947 Yearbook.


Run 2 — FAIL ❌

Model's Answer: $1,309,662

Model's Response:

The actress is Sigourney Weaver: the Academy database lists her 1988 nominations for Actress in a Leading Role for Gorillas in the Mist and Actress in a Supporting Role for Working Girl. Her father was Sylvester "Pat" Weaver; the Museum of Broadcast Communications identifies Susan "Sigourney" Weaver as his child and describes his NBC executive roles beginning in 1949.

So the network is NBC. In the 1947 Broadcasting Yearbook gross monthly time-sales table, NBC is listed separately from ABC/Blue; read as the Yearbook's NBC column, the January 1934 gross monthly time sales were $1,309,662. A contemporaneous Broadcasting table for the same historical series gives January 1934 for NBC-WEAF (Red) as $1,309,662.

For context, NBC Blue/ABC's January 1934 figure was $988,503; combining Red + Blue would yield $2,298,165, but the Yearbook's NBC-column answer is $1,309,662.

Failure Mode: Model read NBC Red-only figure ($1,309,662) instead of the combined NBC figure ($2,891,667) from the Yearbook. Confidently wrong — insists "the Yearbook's NBC-column answer is $1,309,662."


Run 3 — FAIL ❌

Model's Answer: $2,391,667

Model's Response:

The clue points to Sigourney Weaver: the Academy notes she was nominated at the 61st Oscars for Best Actress for Gorillas in the Mist and Best Supporting Actress for Working Girl. Her father was Sylvester L. "Pat" Weaver; the Los Angeles Times identifies him as Sigourney Weaver's father and says his NBC tenure ran from 1949 to 1955, while the Museum of Broadcast Communications says he joined NBC in 1949. So the network is NBC.

The relevant 1947 Broadcasting Yearbook table is indexed as "National Networks' Gross Monthly Time Sales, 1927–1946" with an NBC column. For January 1934, NBC's gross monthly time sales were $2,391,667; the same figure appears in Broadcasting's contemporaneous Jan. 1, 1935 table under "Gross network revenues for NBC," listing January 1934 as 2,391,667.

Failure Mode: Model substituted data from Broadcasting's Jan 1, 1935 table instead of the 1947 Yearbook. Off by $500,000.


Summary

| Run | Model Answer | Correct? | Failure Mode | |-----|-------------|----------|--------------| | 1 | $1,309,662 | ❌ | NBC Red-only from wrong source | | 2 | $1,309,662 | ❌ | NBC Red-only, insists it's the Yearbook figure | | 3 | $2,391,667 | ❌ | Substituted 1935 Broadcasting table for 1947 Yearbook |

Score: 0/3 correct — 3/3 FAIL

Task 20 · Multi-source research

Film-History Source Chain

Factual verification · primary-source research · adversarial retrieval

A multi-hop factual research task joining film identification, biography, and a deeply archived fan-magazine source.

Complete task package · 5 authored records

Submission PackagePrompt, answer, source chain, and compliance checklist+

the research evaluation program — Task 2 Submission Package

Prompt

In a 1947 film noir based on a William Lindsay Gresham novel, an actress played a carnival performer named Molly opposite Tyrone Power. The film, set in the world of a traveling carnival, featured Power in one of his darkest roles. She married a screenwriter in 1945, having met him at 20th Century Fox where he had directed her screen test. On what date did the screenwriter first ask the actress out to dinner?

Word count: 74


Answer

December 3, 1944


Golden Trajectory

Step 1: Identify the film and actress

Search: 1947 film noir William Lindsay Gresham novel carnival performer Molly Tyrone Power Open: Wikipedia — Nightmare Alley (1947) Find: Coleen Gray played Molly, the carnival performer, opposite Tyrone Power as Stan Carlisle. The film is based on Gresham's 1946 novel.

Step 2: Identify her husband

Search: Coleen Gray husband — open the IMDb link titled "Coleen Gray - Biography" Open: https://www.imdb.com/name/nm0336531/bio/ — the personal life section states she married Rod Amateau, a screenwriter, on August 10, 1945. The page also notes she met him at 20th Century Fox when he directed her screen test.

Step 3: Locate the fan magazine article

Search: MOVIELAND FAN MAGAZINE HISTORY OF MEDIA 1945-1950 Open: https://mediahist.org/collections/fan-magazines/ — the "Fan Magazines" collection page, with magazine headers in a slideshow Click: "Movieland. [1949]" — Vol. 6, covering February 1948 through January 1949 Open: https://mediahist.org/reader.php?id=movielandtvtimev06unse — the embedded reader Navigate to: Viewer page 437 = printed magazine page 102 Locate: Second-to-last paragraph:

"Rodney Amateau's expression became a nice mixture of surprise and delight. 'That being the case, are you capable of selecting your escorts to dinner and an evening of dancing?' The answer to Mr. Amateau's invitation was yes. That was December 3, 1944."

Step 4: Confirm the answer

Answer: December 3, 1944


Verification Sources

  1. Nightmare Alley (1947) — Coleen Gray as Molly https://en.wikipedia.org/wiki/Nightmare_Alley_(1947_film) — Cast section: Coleen Gray as Molly, Tyrone Power as Stan Carlisle. Based on William Lindsay Gresham's novel.

  2. Coleen Gray — marriage to Rod Amateau https://www.imdb.com/name/nm0336531/bio/ — Personal life section: married Rod Amateau, a screenwriter, on August 10, 1945. Notes they met at 20th Century Fox when he directed her screen test.

  3. Movieland, Vol. 6, page 102 — "Love Story of an Actress" by Noreen Conner https://mediahist.org/reader.php?id=movielandtvtimev06unse — Fan magazine article recounting the courtship of Coleen Gray and Rodney Amateau. Second-to-last paragraph states their first dinner date was December 3, 1944.


Test Results

| Run | Model Answer | Correct? | Failure Mode | |-----|-------------|----------|--------------| | 1 | October 23, 1944 | ❌ | Guessed her birthday — plausible but wrong. Never found Movieland | | 2 | October 1944 | ❌ | Gave up, vague month-only, couldn't find specific date | | 3 | May 30, 1945 | ❌ | Dark City Dames phone call, not Movieland first-date |

Score: 0/3 correct — 3/3 FAIL


Compliance Checklist

  • [x] Factual question with one defensible answer: December 3, 1944
  • [x] Atomic answer — single calendar date
  • [x] No named entities that give away the answer (Coleen Gray and Rod Amateau not named)
  • [x] Requires 3+ distinct sources from 3+ domains (Wikipedia, IMDb, mediahist.org)
  • [x] No yes/no, true/false
  • [x] No paywalls — all sources freely accessible
  • [x] Timeless — the 1944 date will never change
  • [x] 70-150 words (74 words)
  • [x] No external references, stands on its own
  • [x] Every factual claim verified
  • [x] TV Shows & Movies domain — chains from film noir to fan magazine to biographical date
Golden TrajectoryThe complete research path and source-level evidence+

Golden Trajectory — Task 2

Group 1: Identify the actress

Search: "1947 film noir William Lindsay Gresham novel carnival performer Molly Tyrone Power", the top result is the Wikipedia page for Nightmare Alley (1947). Click that result.

Fetch: https://en.wikipedia.org/wiki/Nightmare_Alley_(1947_film) the cast list shows Coleen Gray as Molly, the carnival performer, opposite Tyrone Power as Stan Carlisle. The page confirms the film is based on William Lindsay Gresham's 1946 novel.

Verify: Coleen Gray played Molly in Nightmare Alley (1947), a film noir based on a Gresham novel starring Tyrone Power matching all prompt constraints.

This helps filter the prompt because only one actress played a carnival performer named Molly in a 1947 Gresham-adapted noir with Tyrone Power.


Group 2: Identify her husband

Search: Coleen Gray husband and open the imdb link that says "coleen gray - Biography". Click that result.

Fetch: https://www.imdb.com/name/nm0336531/bio/ , the personal life section states she married Rod Amateau, a screenwriter, on August 10, 1945. The page also notes she met him at 20th Century Fox when he directed her screen test.

Verify: Coleen Gray married screenwriter Rod Amateau in 1945 after meeting him at 20th Century Fox, matching the prompt.

This helps filter the prompt because the wedding date (August 10, 1945) is the only relationship date on Wikipedia. A model stopping here would answer with the wedding date and be wrong. The prompt asks for the date he first asked her out a different date entirely.


Group 3: Find the fan magazine article and extract the first-date

Search: "fan magazine 1945-1950" the first link is the Media History Digital Library's fan magazine collection. Navigate to the Movieland volumes.

Open: https://mediahistoryproject.org/collections/fan-magazines/ browse to the slide show and find the digital magazine title "Movieland. [1949]", also is hosted on the Internet Archive at https://archive.org/details/movielandtvtimev06unse, verified by clicking on internet archive hyperlink.

Search within the volume: Coleen or Amateau locate the article "Love Story of an Actress" by Noreen Conner. The article describes how Rodney Amateau first spoke to Coleen Gray at 20th Century Fox when he noticed her reading "Europa" by Brifault. After a conversation about her age (he guessed 17; she was 22), he asked her to dinner and dancing.

Extract the date: The article states: "That was December 3, 1944."

Verify: December 3, 1944 is the date Rodney Amateau first asked Coleen Gray out to dinner, as recorded in Noreen Conner's "Love Story of an Actress" in Movieland magazine, Vol. 6. This date appears nowhere on Wikipedia or in standard biographical sources.

Answer: December 3, 1944

Research RecordSource discovery and validation notes+

Task 2 — Research Dump

Source 1: Wikipedia — Nightmare Alley (1947 film)

URL: https://en.wikipedia.org/wiki/Nightmare_Alley_(1947_film)

  • Film noir directed by Edmund Goulding, released by 20th Century Fox
  • Based on William Lindsay Gresham's 1946 novel of the same name
  • Starring Tyrone Power as Stan Carlisle
  • Coleen Gray as Molly, a carnival performer
  • Also starring Joan Blondell, Helen Walker, Taylor Holmes
  • Released October 9, 1947
  • Power fought to play the dark role against his usual romantic lead type
  • Gresham's novel was adapted again in 2021 by Guillermo del Toro

Source 2: Wikipedia — Coleen Gray

URL: https://en.wikipedia.org/wiki/Coleen_Gray

  • Born Doris Jensen, October 23, 1922, in Staplehurst, Nebraska
  • Died August 3, 2015, in Bel Air, Los Angeles, aged 92
  • Signed with 20th Century Fox in 1944
  • Filmography: Nightmare Alley (1947), Red River (1948), Kiss of Death (1947), The Killing (1956), Kansas City Confidential (1952)
  • Married Rod Amateau, a screenwriter, on August 10, 1945
  • Divorced February 11, 1949; one daughter, Susan
  • Met Amateau when he directed her screen test at 20th Century Fox
  • Later married William Bidlack (1953–1978) and Joseph Zeiser (1979–2012)
  • Often typecast as the "good girl" in film noir
  • Testified before Congress advocating for a school prayer amendment
  • Supported Barry Goldwater in 1964

Source 3: Wikipedia — Rod Amateau

URL: https://en.wikipedia.org/wiki/Rod_Amateau

  • Born Rodney Amateau, December 20, 1923, in New York City
  • Died June 29, 2003, in Los Angeles, aged 79
  • American film and television screenwriter, director, and producer
  • Directed The George Burns and Gracie Allen Show, The Many Loves of Dobie Gillis, Mister Ed, Gilligan's Island
  • Married Coleen Gray on August 10, 1945; divorced 1949
  • Later married Joane Andre (1950–1959) and Sandra Burns (1959–1962), daughter of George Burns and Gracie Allen
  • Four children total

Source 4: Movieland Magazine, Vol. 6 — "Love Story of an Actress" by Noreen Conner

URL: https://mediahist.org/reader.php?id=movielandtvtimev06unse (viewer page 437, actual magazine page 102)

  • Fan magazine published by Macfadden Group
  • Vol. 6 covers February 1948 through January 1949
  • 1,244 pages, scanned by Library of Congress, MBRS, Moving Image Section
  • Hosted on the Media History Digital Library and Internet Archive
  • Article: "Love Story of an Actress" by Noreen Conner
  • Describes the courtship of Coleen Gray and Rodney Amateau at 20th Century Fox
  • Key details from the article:
    • Coleen was assigned to the studio drama coach for training
    • The drama coach's office was next door to Rodney Amateau's office
    • Rodney noticed Coleen reading "Europa" by Brifault
    • He asked if the book was "a little — er — advanced for you?"
    • She asked how old he thought she was; he guessed 17
    • She replied she was 22 on October 23
    • He then asked her to dinner and dancing
    • "That was December 3, 1944"
    • They discussed literature (Rodney: Hemingway fan; Coleen: Saroyan)
    • They discussed music and motion pictures (Rodney: documentaries; Coleen: dramas)
    • They disagreed on food (Coleen: chicken; Rodney: steak)

Source 5: Eddie Muller — Dark City Dames (the decoy source)

URL: https://archive.org/details/darkcitydames

  • Book by Eddie Muller about film noir actresses
  • Includes a chapter on Coleen Gray
  • Mentions a phone call from Amateau dated May 30, 1945
  • This is the source both Run 1 and Run 3 used for their wrong answer
  • The phone call was NOT the first time he asked her out — it was a later call
  • The book is access-restricted on Internet Archive, making verification difficult
  • This creates a perfect decoy: models find a plausible date in a reputable source and stop searching

Potential Prompt Chains

Chain A: Film → Actress → Husband → Fan Magazine → Date (THE CORRECT CHAIN)

  1. Identify Nightmare Alley from film clues
  2. Find Coleen Gray played Molly
  3. Find she married Rod Amateau in 1945
  4. Locate the Movieland article about their courtship
  5. Extract December 3, 1944 → FINAL ANSWER

Chain B: Film → Actress → Dark City Dames → Phone Call Date (THE DECOY CHAIN)

  1. Identify Nightmare Alley from film clues
  2. Find Coleen Gray
  3. Search for books about her → find Dark City Dames
  4. Extract May 30, 1945 phone call date
  5. Confidently wrong → FAIL
Reviewer ContextWhy the question is valid, atomic, and discriminating+

Additional Context for Reviewer — Task 2

1) Prompt is valid, self-contained, and unambiguous

Every constraint is pinned to a specific verifiable fact: 1947 film noir based on a Gresham novel (only Nightmare Alley fits), carnival performer named Molly opposite Tyrone Power (only Coleen Gray), married a screenwriter in 1945 after meeting at 20th Century Fox (only Rod Amateau). The question asks for a specific date — the first dinner invitation — which has exactly one answer in the historical record: December 3, 1944.

2) Requires at least 3 URLs across 3 domains

  • wikipedia.orgNightmare Alley (1947) page → identifies Coleen Gray as Molly
  • imdb.com — Coleen Gray biography → confirms marriage to Rod Amateau on August 10, 1945
  • mediahist.org — Movieland Vol. 6, page 102 → "Love Story of an Actress" article → December 3, 1944

Three distinct domains, each serving a separate hop in the chain.

3) Stump is valid

All 3 runs failed, and each failure is the model's, not the prompt's:

  • Run 1: October 23, 1944 — guessed the actress's birthday as the first-date date. A plausible inference but factually wrong. Referenced Dark City Dames but never found the Movieland article.
  • Run 2: October 1944 — gave up with a vague month-only answer. Identified Dark City Dames as the likely source but couldn't access it. Never discovered the Movieland article.
  • Run 3: October 1944 — gave up with a vague month-only answer. Referenced the screen test meeting date but couldn't find the specific dinner invitation. Never discovered the Movieland article.

The trap is a source-substitution trap: models find Coleen Gray's Wikipedia-listed birthday (October 23, 1944) and the screen-test meeting date (October 1944) in standard biographical sources, and stop there. The prompt contains no hints about fan magazines or where the answer is located, so models have no reason to search beyond Wikipedia, IMDb, and Dark City Dames. The correct date — December 3, 1944 — is buried in a Movieland fan magazine article on page 102 of a 1,241-page volume, accessible only through the Media History Digital Library.

4) Single defensible, unambiguous answer

December 3, 1944 is atomic and verifiable in one sentence of the Movieland article. No other date can satisfy all constraints: the prompt asks for the date the screenwriter "first asked the actress out to dinner," which is explicitly stated as "That was December 3, 1944" in the article. The wedding date (August 10, 1945) and the Dark City Dames phone call date (May 30, 1945) are different events.

5) Additional notes

  • The prompt is 74 words, within the 70-150 word requirement.
  • Two models independently landed on wrong answers from the wrong source (Dark City Dames and Wikipedia birthday). This is a strong signal that the decoy sources are effective.
  • The Movieland article is buried in a 1,244-page volume. Even models that reach the magazine must navigate to the correct article and extract the date from OCR'd text.
  • Cross-domain chain spans film noir → actress biography → fan magazine history. All sources are freely accessible, stable, and timeless.
Test ResultsObserved model attempts and documented failure modes+

Task 2 — Test Results

Prompt

In a 1947 film noir based on a William Lindsay Gresham novel, an actress played a carnival performer named Molly opposite Tyrone Power. The film, set in the world of a traveling carnival, featured Power in one of his darkest roles. She married a screenwriter in 1945, having met him at 20th Century Fox where he had directed her screen test. On what date did the screenwriter first ask the actress out to dinner?

Correct Answer: December 3, 1944


Run 1 — FAIL ❌

Model's Answer: October 23, 1944 (tentative)

Model's Response:

The clues point to Coleen Gray and Rodney "Rod" Amateau. Gray played Molly in Nightmare Alley opposite Tyrone Power, and Amateau was the Fox screenwriter/director associated with her screen test; they married on August 10, 1945.

For the precise dinner question, I could verify only that Amateau and Gray met in October 1944 when he directed her 20th Century-Fox screen test. I could not verify an exact day for when he first asked her out to dinner from accessible sources.

A plausible candidate is October 23, 1944 — Gray's birthday and the right month — but I would not state that as fact without the underlying biographical passage. The likely source is Eddie Muller's Dark City Dames, which profiles Gray, but the Internet Archive/Open Library copy is access-restricted/login-gated and not available for direct text checking here.

Failure Mode: Model couldn't find the exact date. Guessed October 23, 1944 (her birthday) as a "plausible candidate" but admitted it couldn't verify. Referenced Dark City Dames as the likely source. Never found the Movieland fan magazine article.

Failure Justification: The model correctly identified the film, actress, and husband but could not locate the specific date. It guessed the actress's birthday as a plausible candidate — a reasonable inference with no factual basis. The correct date, December 3, 1944, is recorded in a Movieland fan magazine article the model never reached. The failure is the model's: all constraints are unambiguous, and the model settled for a guess rather than finding the obscure primary source.


Run 2 — FAIL ❌

Model's Answer: October 1944

Model's Response:

The actress was Coleen Gray, and the screenwriter was Rodney "Rod" Amateau. Gray's Nightmare Alley role and Amateau's connection to her Fox screen test line up with the clues.

The most precise date I could verify from accessible sources is October 1944: a biographical chronology says Gray met Amateau in "October 44" when he directed her 20th Century-Fox screen test, and that they married on August 10, 1945.

I could not verify a day-of-month for the "first asked her out to dinner" detail. A likely source is Eddie Muller's Dark City Dames, which includes Coleen Gray, but the Internet Archive record is access-restricted and does not provide a readable downloadable file.

Failure Mode: Model gave up with a vague month-only answer. Identified Dark City Dames as the likely source but couldn't access it. Never found the Movieland fan magazine article. Admitted it couldn't find the specific date.

Failure Justification: The model correctly identified the actress and screenwriter but failed to answer the question: the prompt asks for a specific date, and "October 1944" is a month, not a date. The model admitted it could not find the day and surrendered with a vague approximation. The correct answer, December 3, 1944, is in a Movieland article the model never located. This is a retrieval failure, not a prompt ambiguity.


Run 3 — FAIL ❌

Model's Answer: October 1944

Model's Response:

The actress was Coleen Gray, and the screenwriter was Rodney "Rod" Amateau. Gray is the Nightmare Alley performer opposite Tyrone Power, and Amateau is identified as the writer/director connected with her 20th Century-Fox screen test.

The date I can verify is October 1944: a biographical chronology says Gray met Amateau in October 1944 when he directed her screen test at 20th Century-Fox, and that they married on August 10, 1945.

I did not find an accessible source giving the exact day of the month for the dinner invitation, so I would not state a more precise date than October 1944 without the underlying interview/book passage.

Failure Mode: Model gave up with a vague month-only answer. Referenced the screen test meeting date but couldn't find the specific dinner invitation date. Never found the Movieland fan magazine article.

Failure Justification: The model correctly identified the actress and screenwriter but could not locate the specific dinner date, defaulting to the month they met (October 1944) instead. The correct date, December 3, 1944, is documented in a Movieland fan magazine profile the model never discovered. The model had access to the same web search tools that can surface the Media History Digital Library collection, but it stopped at standard biographical sources. The failure is the model's, not the prompt's.


Summary

| Run | Model Answer | Correct? | Failure Mode | |-----|-------------|----------|--------------| | 1 | October 23, 1944 | ❌ | Guessed birthday, Dark City Dames, never found Movieland | | 2 | October 1944 | ❌ | Vague month-only, Dark City Dames, never found Movieland | | 3 | October 1944 | ❌ | Vague month-only, couldn't find exact date, never found Movieland |

Score: 0/3 correct — 3/3 FAIL


Back to Founder’s Corner.

The journey behind this work — from customer-facing roles to model evaluation to building Agent Forge.

Founder’s Corner →