Connected
Investing.comSharkNinja Turns to Palantir and AWS to Optimize Promotions and Media· Business WireSharkNinja Puts Part of a $247M Tariff Refund Back Into AI· ZacksSharkNinja's Jailbreak Program Reaches 400 Departmental AI Projects· SalesforceHow SharkNinja Turned Fragmented Data Into Ecommerce Momentum· Business InsiderWhy SharkNinja Paused the Whole Company for a Four-Day AI Bet· SalesforceHow SharkNinja Turned Product Unboxing Into a Guided AI Conversation· The Motley FoolSharkNinja Commits to an AI-First Operating Model in Q1 2026 Results· VentureFizzJailbreak LIVE: SharkNinja's Four-Day Global AI Hackathon· Fast CompanySharkNinja Launches $1M AI Challenge for All 4,200 Employees· SharkNinjaSharkNinja Builds an AI-Native Early Talent Program· ZacksSharkNinja's AI Roadmap Signals Long-Term Upside for the Stock· Inc.SharkNinja Named to Inc.'s Best in Business 2025 for Innovation and Marketing· Business WireSharkNinja and Boston University Launch Dedicated AI & Analytics Lab· SiliconANGLEHow AI Is Transforming the Consumer Experience at SharkNinja· Digital Commerce 360SharkNinja Adds AI Agent to Unified DTC Ecommerce Site· TIMESharkNinja Named to TIME100 Most Influential Companies of 2025· Fast CompanySharkNinja Named to Fast Company's World's 50 Most Innovative Companies of 2025·

August 2026 · 8 min read

The Grunt Work Was the Apprenticeship

Judgment used to come free, bolted onto the boring work. Agents just automated the boring work. Here's how I'd rebuild the apprenticeship on purpose, before a generation of juniors never learns to catch a confident wrong answer.

Judgment is a byproduct. You get it doing dull work at a desk next to someone who’s already good, and nobody ever writes it down as training.

When I came up as an analytics developer, every data mart shipped with a tax. Before anyone would trust it, I had to prove it was right, and there was no button for that. So I wrote the data-quality checks by hand, dozens per table. Row counts against the source. Summary stats that had to reconcile down to the cent. Most of them were boring. A couple were the only thing standing between me and a number that was quietly wrong.

Early on I built a revenue mart, ran my checks, and the row count matched the source exactly. I called it done. My lead asked me one thing: what’s the total come to? It came in about fifteen percent high. A join to a dimension had fanned out and doubled a slice of the rows. The count still matched. The money didn’t.

Here’s what I took from that, and I still use it. Matching the source proves nothing. It proves you copied the source, bugs and all. The check that counts ties your number to one somebody outside the team already trusts, the figure finance closed the month on, the total the board already saw. You reconcile to the thing you can’t quietly fake.

That’s where the judgment came from. Not a class. A desk, a pile of throwaway queries, and being wrong in small ways until I could feel a bad number before I could prove it. It rode along under the paid work, free, and nobody put it on a plan.

Agents just automated the desk.

Point Cursor or Claude Code at that pile of tedium and it’s gone by lunch. Everyone reads that as a clean win. It isn’t. The dull work was never in the way of the learning. It was the learning.

That’s the whole mechanism, and the argument turns on it. You learn to judge by committing to an answer and then colliding with a truer one. The revenue mart was one collision. A few hundred more and you stop needing them, because you can feel a wrong answer before you can prove it. The grunt work was just the cheapest way anyone ever found to run the collision that often.

The early data already points this way. Stanford’s Digital Economy Lab found employment for 22-to-25-year-olds in the most AI-exposed jobs dropping over the past year, while it held steady for older workers doing the same work. That measures hiring, not judgment. Jobs move first. But the jobs thinning out are the exact ones people used to learn on.

None of this breaks this quarter, which is why nobody stops it. You automate the junior’s work and the team gets cheaper by Friday. The cost shows up eight years later, when your seniors retire and there’s nobody behind them who ever climbed. Judgment takes a decade to grow. You can stop growing it this afternoon and not feel it until the decade’s up. Every field that automated a skill this way paid the same bill. Aviation ran the experiment first with autopilot, and the dependency stayed invisible right until the day the automation quit and the hands weren’t there.

So build the collision back on purpose. Here’s what I’d actually do.

Put the junior between two things. Below them sits the agent, doing the hands. Above them sits a senior, out of the execution loop now because the machine is in it. The junior owns the one piece neither can hold for them: deciding whether the output is right, and whether it ships.

The two-pairing apprenticeship loopThe junior predicts a task before prompting. A coding agent executes it and makes a decision. A separate compare skill takes the junior's prediction and the coding agent's decision and drafts the record. The junior judges the gap in that record; the record and the junior's read go to the senior and junior to review together, who calibrate. The junior decides: recoverable work ships, high-cost work escalates. Each cycle the line moves out and judgment compounds.The two-pairing loopThe junior predicts, compares, decides. A coding agent executes; a compare skill scribes the record.SENIOR · PAIR UPReviews it with youreads your reasoning,calibrates the callyour readcalibratea taskJUNIOR · PREDICTSPredictcommit the answerand risks, firstJUNIOR · COMPARESComparejudge the gap inthe recordJUNIOR · DECIDESDecideown it and ship,or escalatethe gaprecoverableshipCODING AGENT · PAIR DOWNExecutesdoes the hands,makes a decisionCOMPARE SKILLscribes the record;surfaces, doesn't judgerunpredictionits decisionthe recordeach cycle the line moves out · scope widens, judgment compounds
The loop, once you build it back on purpose. The junior guesses before the agent runs, and the gap between the guess and the output is what the senior reads.

The move that carries it is small. Make the junior answer first. Before they open Cursor, they commit to a result. What it should return, and where it’s most likely to break. Then they run the agent and compare. The gap between the guess and the output is the collision, staged on purpose, and it’s sharper than the old version because the disagreement is sitting right there on the screen. Prompt first and think second, and you’ve skipped the rep. Guess first, and you’ve turned the agent into a sparring partner that hands you a hundred reps a week where the old desk gave you five.

The senior’s job changes too. Don’t have them re-check the agent’s output. That turns your best person into a bottleneck, and it will not survive a real quarter. Have them read the junior’s reasoning instead. What you expected, and where you and the machine parted ways. That’s a five-minute read, not a redo. And it’s time they’re spending anyway, cleaning up bad output after it ships. You’re just moving it earlier, to where it teaches.

A small tool helps. Call it a compare skill. It sets the junior’s guess next to the agent’s output and marks every place they disagree. The comparison gets written down, and there it stops. Settling the disagreement is the rep, and the rep belongs to the junior. Nobody acts on what the skill says until a person decides to, which is the whole line between a skill and an agent. (Run it as a separate pass from whatever wrote the code. An agent grading its own homework grades gently.) What comes out is one page. The senior reads the page. That’s the apprenticeship now. A page, and five minutes.

Don’t put a senior on every call. Ration by what being wrong costs. Recoverable work the junior owns and reports on afterward, and yes, some of it ships wrong. That’s tuition, and it’s cheap. The stuff you can’t take back, the credit that goes out to accounts that don’t qualify, the number that lands in a board deck, that gets a senior before it ships. Teaching a junior to tell those two apart is most of the job. Someone who escalates everything is as useless as someone who escalates nothing.

None of it happens unless it’s on the plan with hours against it. Growth is a line item or it’s nothing. If the only time a junior sits with a senior is when a deadline forces it, you don’t have an apprenticeship. You have someone who gets watched by accident, once in a while.

Here’s the part people won’t like. If you can’t staff the seat, don’t make the hire. An intern parked in front of Claude with nobody reading their reasoning is a liability you’re training to rubber-stamp, and in three years that liability is your senior. Fix the staffing or skip the hire. Don’t take the intern and hope.

This only works where the work is checkable. If a junior can predict the answer and check it, the loop runs. For the genuinely open problem, the one where nobody knows the answer walking in, there’s no guess to commit and no gap to read, and that work still needs a senior sitting inside it the old expensive way. The loop doesn’t rescue you there. It buys back the hours so you can afford to be there when it counts.

Seniority used to come free. It rode along under the boring work, and nobody put it on a budget, so nobody noticed it was a subsidy. The subsidy is over. The boring work still gets done, faster than it ever has, and the judgment that used to come attached to it does not come anymore. You build the desk back on purpose, the junior between a senior and a machine with one page to argue over, or you graduate a generation of juniors who never learned to catch a confident wrong answer. Right when the machines start producing them faster than anyone can check.

  • leadership
  • org-design
  • ai
  • data-strategy

All insights