The Greenbox Story · Losing the Thread

Outcome Over Output: The Feature Factory

· 43 min read

Charlotte built the discovery cadence at the start of the year, while Brisbane was still in pilot. She’d been reading Teresa Torres on continuous discovery habits, and she adapted the core of it to Greenbox’s three-squad shape: talk to subscribers every week, check your assumptions, connect everything to outcomes.

The rhythm fitted on an index card. Monday, an assumptions check, five minutes added to standup: “What’s the riskiest assumption we’re carrying right now?” Most weeks the answer was “nothing new”; the point was the habit. Tuesday, one subscriber interview per squad, fifteen to twenty minutes, rotating interviewer, with the LLMA neural network trained to predict the next token in a sequence, large enough that it generalises to tasks it wasn’t explicitly trained for. transcribing and posting summaries to a channel all three squads could see. Wednesday, Example Mapping for the stories that needed it. Thursday and Friday: build, ship, measure. Around the weekly loop, slower wheels: a fortnightly retro, a monthly Impact Map review, a quarterly look at the Wardley Map.

It paid off within weeks. A routine Tuesday interview with a Brisbane prospect confirmed what the pilot data only hinted at: single-person households wanted a smaller, cheaper box. The squad pivoted the launch plan, and the small box outsold the original lineup three to one. Without that conversation, the team would have concluded that Brisbane didn’t want produce boxes, when the real answer was that Brisbane wanted different produce boxes.

It caught quieter things too. Sam’s Tuesday interview with Mrs Patterson, subscribed since the very first box, ended with a line Sam screenshotted and posted to the team channel without commentary: “I’ve never met any of you, but I feel like you know me.” Nobody replied for twenty minutes. Then Priya reacted with a thumbs-up. Then Maya. Then Tom. One by one, every person in the company.

That was three months ago.

The quarterly review

The quarterly review happens on a Thursday in late April. Three squads present. The format is simple: what did you ship, what’s the impact, what’s next.

Perth goes first. Tom runs through the slide. “We shipped the new box preview redesign, the subscription gifting feature, the improved farm dashboard, and the allergen warning system. Four features. All on time.”

Melbourne: Anika presents. “Meal kit expansion phase two, the corporate ordering portal, delivery window preferences, and the subscriber referral programme. Four features.”

Brisbane: “The new small-box pricing tier, farmer spotlight pages, a Brisbane-specific landing page, and push notification preferences. Four features.”

Twelve features in one quarter. The team applauds. Maya smiles. It feels like progress.

Charlotte doesn’t applaud. She’s writing something in her notebook.

After the presentations, she asks one question.

“Which of these twelve features is responsible for your subscriber growth this quarter?”

Silence.

Tom tries first. “The box preview probably helped conversion. We saw a bump in signups after it launched.”

“How big a bump?”

“I’d have to check the numbers.”

“Did you set a target before you built it?”

Tom pauses. “Not… specifically. We knew it would help.”

Charlotte turns to Anika. “The corporate portal. How many corporate orders have come through?”

Anika checks her laptop. “Seventeen.”

“What was the target?”

“We didn’t set one. The request came from three enterprise prospects.”

Charlotte looks at the Brisbane squad. “The farmer spotlight pages. What metric do they affect?”

“Engagement. People like reading about the farmers.”

“How do you know?”

“We’ve had positive feedback.”

“How many people read them?”

Another pause. “We haven’t added analytics to those pages yet.”

Charlotte closes her notebook. “You shipped twelve features. You can’t measure the impact of any of them. You’re a feature factory.”

The mirror

The term stings because it’s accurate.

A feature factory is a team that measures productivity by output, features shipped, stories completed, velocity maintained, without connecting that output to outcomes. The features get built. They get deployed. They get announced. Nobody checks whether they changed anything.

Melissa Perri coined the term “build trap” for this pattern: staying busy building the wrong things. Not wrong in the sense that they’re bad features. Wrong in the sense that they might not matter. The team is productive. The team is efficient. The team is potentially wasting its time.

Maya recognises the pattern because she’s part of the cause. Over the last three months, she’s said yes to almost everything. A subscriber asked for gifting. Maya added it to the backlog. An enterprise prospect wanted corporate ordering. Maya flagged it as a priority. A board member mentioned push notifications. Maya asked Brisbane to build it. Each request was reasonable. Each feature was well-built. But the backlog grew by accretion, not by design.

“I kept saying yes,” Maya tells Charlotte after the review. “Every request seemed important when it arrived.”

“They probably were important to the person asking. But important to one subscriber isn’t the same as important to the business. You stopped asking ‘why’ and started asking ‘when.’”

Where the discovery went

Charlotte pulls up the weekly cadence she’d designed three months earlier. Monday assumptions check. Tuesday subscriber interview. Wednesday Example Mapping. Thursday and Friday: build, ship, measure.

“When did the Tuesday interviews stop?”

Tom answers honestly. “About two months ago. We were heads down on the gifting feature. It had a hard deadline for Easter. The interviews felt like they could wait.”

“And the Monday assumptions check?”

“We still do it. But nobody’s raised a risky assumption in weeks. We’ve been doing features we understand.”

“Or features you think you understand.”

Charlotte isn’t angry. She’s seen this before, at three previous companies, in fact. The pattern is always the same. A team builds good discovery habits. The habits work. The team gets confident. Delivery pressure builds. The discovery activities feel like overhead because the team is building things they’ve already decided to build. One week the interview gets skipped. Then another. Then the Example Mapping sessions stop because all the stories are “Clear.” Within two months, the team is back to building from assumptions, the same assumptions that Assumption Mapping and Impact Mapping were designed to catch.

“Discovery didn’t stop because you decided to stop,” Charlotte says. “It stopped because delivery pressure created a gradient, and the team rolled downhill.”

The backlog audit

Charlotte suggests a backlog review. The kind nobody wants to do.

The Greenbox backlog has 308 items. They sit in a project management tool, loosely prioritised, spanning three squads and eighteen months of accumulated requests. Some were added by Maya. Some by Tom. Some by subscribers. Some by board members. A few were added by people who no longer work at Greenbox.

Charlotte asks each squad to go through their items and answer one question: what outcome does this item achieve?

Not “what does it do.” What outcome. What measurable change in a number the business cares about, subscriber growth, churn reduction, revenue per subscriber, delivery cost, NPS score.

The Perth squad reviews their 120 items in ninety minutes. Twice during the session Tom’s laptop chimes with a review request, and twice he approves it with a one-word “LGTM” without scrolling past the first file. When they’re done, he reads the results.

“Forty-seven items have a clear outcome. Thirty-one have a vague outcome, something like ‘improve the subscriber experience.’ Forty-two have no outcome at all. They’re features for their own sake.”

Anika’s Melbourne squad: “Thirty-nine with outcomes. Twenty-eight vague. Fifty-three with no outcome. Three items that nobody can explain the purpose of.”

Brisbane: “Twenty-two with outcomes. Nineteen vague. Twenty-seven with no outcome.”

Charlotte writes the totals on the whiteboard.

108
Clear outcome
(35%)
78
Vague outcome
(25%)
122
No outcome
(40%)

Forty percent of the backlog has no measurable outcome attached. Two items in five. The team has been feeding features into a machine without checking whether anything comes out the other end.

“This is what a feature factory looks like from the inside,” Charlotte says. “Nobody made a bad decision. Each item was added for a reason. But without an outcome, you can’t prioritise, you can’t measure, and you can’t learn.”

The impact map test

Lee joins the conversation by video. He pulls up the latest revision of the impact map, the one the monthly reviews are meant to meant to test. The goal at the root is specific: reduce churn from 6% to 4%.

“Let’s check. Which of the twelve features you shipped this quarter are on this map?”

Tom scans the map. The box preview redesign connects to “improve first-box experience,” which connects to reducing churn. That’s on the map.

The allergen warning system, born from the allergen incident, connects to “prevent bad box experiences,” which also connects to reducing churn. That’s on the map.

The other ten features? Gifting, corporate ordering, delivery windows, referral programme, pricing tiers, farmer spotlights, landing pages, push notifications, farm dashboard, meal kit expansion. None of them appear on the impact map.

“Two out of twelve,” Lee says. “You spent roughly 80% of your development capacity on work that wasn’t connected to your stated goal.”

“Some of those are about growth, not churn,” Maya objects.

“Fair. What’s the growth goal?”

Maya hesitates. “We want to grow.”

“That’s an aspiration, not a goal. How much growth? By when? Through which channels? The impact map is specific. ‘Reduce churn from 6% to 4%.’ You can measure that. You can connect work to it. ‘We want to grow’ lets you justify building anything.”

Output vs outcomes

Charlotte draws two columns on the whiteboard.

Output Outcome
Shipped gifting feature Subscriber growth increased by X%
Shipped corporate portal Revenue from corporate accounts reached $Y/month
Shipped referral programme Z new subscribers joined via referral
Shipped farmer spotlights NPS improved by W points

“The left column is what you celebrate. The right column is what matters. You can fill in every row on the left. You can’t fill in a single row on the right.”

“Because we didn’t set targets,” Priya says quietly. She’s been listening the whole time.

“Because you didn’t set targets before you built. If you’d said ‘the corporate portal needs to generate fifty orders in the first month or we’ll reconsider,’ you’d have known at seventeen that something was wrong. You’d have talked to those corporate prospects again. You’d have learned. Instead, seventeen feels like progress because you didn’t define what success looked like.”

Ravi, who has spent the quarter on loan to whichever squad was furthest behind, speaks up. “I spent three weeks on the referral programme. How many referrals have come through?”

Sam checks. “Eleven.”

Ravi does the maths in his head. Three weeks of a senior developer’s time to generate eleven referrals. At Greenbox’s average subscriber value, the referral programme has cost roughly twenty times more to build than it has generated.

“I could have spent those three weeks reducing churn,” he says. “One percent less churn would have saved more subscribers than eleven referrals.”

“Now you’re thinking in outcomes,” Charlotte says.

The dual-track reset

Charlotte proposes a reset. Not a revolution, a recalibration.

Dual-track agile. Discovery and delivery run in parallel. Every squad does both, every week. The delivery track builds features. The discovery track validates that the features are worth building.

“You had this,” Charlotte reminds them. “The Tuesday interviews. The Monday assumptions check. You let it lapse because delivery felt more urgent. The fix isn’t a new process. It’s recommitting to the one you already had.”

One discovery activity per squad per week. Non-negotiable. It doesn’t have to be a subscriber interview every time. It could be a data review, a competitor analysis, a prototype test, or a five-minute conversation with Sam about what subscribers are complaining about. The format varies. The habit doesn’t.

Every backlog item gets an outcome statement before it enters a sprint. Not after. Before. “As a [subscriber], I want [feature], so that [outcome]” is the minimum. Better: “We believe [feature] will cause [measurable change] and we’ll know within [timeframe].”

“What about items the board requests?” Maya asks. It’s the question she’s been dreading.

“Same rule. If a board member suggests a feature, you add an outcome. If they can’t articulate one, the item sits in the backlog until someone can. Board requests aren’t immune to the laws of prioritisation.”

Maya nods slowly. This is the part where founder instincts collide with following the process. She’s been saying yes because saying yes feels like progress. Every “yes” is a commitment, a promise, a relationship maintained. Saying “yes, and what outcome are we targeting?” feels slower. It is slower, and the slowness is what stops the bad yes.

Outcome-based roadmap. Instead of a roadmap that says “Q1: gifting, corporate portal, referral programme,” the new roadmap says “Q1: reduce churn from 6% to 4%.” The squads decide what to build to achieve that outcome. Maybe it’s a better onboarding flow. Maybe it’s fixing the three delivery complaints that drive 40% of cancellations. Maybe it’s something nobody has thought of yet, which is why the discovery track exists.

The reconnection

The following Tuesday, every squad does a subscriber interview. For Perth, it’s the first one in nine weeks.

Tom interviews a subscriber named David who’s been with Greenbox for fourteen months. Tom asks the standard questions: what’s working, what’s not, would you recommend it.

David’s answer surprises him. “I nearly cancelled last month. The box was fine. The produce was fine. I just felt like nobody was listening any more.”

“What do you mean?”

“I filled out the feedback form three times about getting too many root vegetables in winter. Nothing changed. I used to feel like someone was paying attention. Now it feels automated.”

Tom thanks him. Ends the call. Sits at his desk.

Three weeks ago, Tom shipped the box preview redesign. It was polished. The code was excellent. David didn’t mention it. What David mentioned was that nobody responded to his feedback. That’s not a feature problem. That’s a relationship problem. And it’s the kind of thing you only learn by talking to subscribers.

Tom posts the summary in the team channel. No commentary. Just the transcript excerpt.

Priya replies first: “This is what we miss when we stop listening.”

The three-week check

Three weeks into the reset, Charlotte checks in.

The Tuesday interviews are happening again. Each squad has conducted three. Two of the nine surfaced insights that changed sprint priorities, a delivery timing complaint in Melbourne that was more widespread than the data showed, and a Brisbane subscriber who explained why she’d downgraded her box size in a way that suggested the pricing tier wasn’t the problem.

The backlog has shrunk from 308 items to 235. Not because they deleted items, because they archived the 73 that nobody could attach an outcome to. Those items aren’t gone. They’re in a “needs outcome” holding area. If someone can articulate why they matter, they come back.

The Monday assumptions check produced one genuinely risky assumption: “We assume corporate customers will reorder monthly.” Nobody had tested this. The Melbourne squad designed a three-email follow-up sequence and tracked reorder rates. The assumption was wrong, corporate customers reorder quarterly, not monthly. The revenue model for the corporate portal was off by 3x.

“That one insight,” Charlotte tells Maya, “is worth more than the last three features you shipped. Because it changes a decision. The features just added code.”

A cafe in Subiaco

On a Wednesday afternoon in late May, Lee and Charlotte meet for coffee at a cafe in Subiaco. They haven’t sat down together properly in months. Lee drove up from Margaret River that morning, his surfboard strapped to the roof.

Charlotte stirs her flat white. “When you first came in, what did you think the problem was?”

“Speed without understanding. The LLMs removed the natural friction that used to force conversations. Implementation became so fast that people stopped talking.”

“And now?”

“Now the team talks first and builds second. But the challenge has shifted. At five people, shared understanding happens naturally. At twenty-eight across three cities, it has to be engineered. And re-engineered, apparently, every time delivery pressure builds.”

Charlotte is honest about how the reset feels. “Some weeks the cadence works beautifully. Other weeks someone skips the interview, the retro gets cancelled.” She looks at her coffee. “I coached a meal kit company before Greenbox. Good people, good product. They went under. I keep checking for the same patterns. Sometimes I push too hard on process because I think if the process is right, the outcome is guaranteed.” She looks up. “It’s not.”

“That’s not a flaw,” Lee says. “That’s experience with scar tissue.”

“What about you? You’ve been pulling back.”

Lee takes a long time to answer. “My ex-wife, Mei, she said I was always coaching other people’s lives. Twenty years of consulting. I’d fly into a company, help them see what they couldn’t see, and fly out. I was doing the same thing at home.” He turns his coffee cup. “Greenbox is the closest I’ve come to building something since I stopped trying to build a marriage.”

“How’s your daughter?”

Lee looks up, surprised. “She’s at university. Environmental science, in Sydney. I’ve started calling her every Sunday.” A half-smile. “She doesn’t always answer. When she does, we talk about carbon sequestration in coastal wetlands.”

His phone buzzes, as if summoned. Yuki: Dad, did you know mangroves sequester carbon 4x faster than terrestrial forests? Lee types back: I did not. Tell me more on Sunday.

Outside, a delivery van with the Greenbox logo pulls up. Charlotte watches it. “That’s Liam, one of the Perth drivers, doing the Subiaco run. Two hundred and forty boxes every Thursday.”

Lee watches the van too. The consultant’s distance gives way to something more personal. “That’s something,” he says. “That’s actually something.”

The market

On a Saturday morning a few weeks after the reset, Maya drives down to the Margaret River farmers’ market. She goes every few weeks, partly for produce, partly because the market is where she first met Dave Morrison, and partly because the three-hour drive through the jarrah forest is the only time she’s unreachable.

Dave is at his usual stall, between the honey seller and the woman who makes goat’s cheese. His son Ben handles the Greenbox supply now, but Dave still comes because he’s been coming since before Ben was born.

They get coffee from the van at the end of the row and sit on an upturned crate behind Dave’s stall. Margaret River cold, the kind that sits in your bones until the sun gets above the trees.

Dave tells her about the frost of 2019. The full story, not the fragments she’s heard before. He woke at 4am to find ice on the inside of the greenhouse plastic. By dawn the entire tomato crop was gone. Three months of work. He didn’t tell Helen until that evening because he spent the day walking the rows, pulling up dead plants, looking for something salvageable. There was nothing.

“I didn’t call anyone. Didn’t tell the other farmers. Just kept going. Replanted the next week.”

Maya is quiet for a while. Then she tells him something she’s never told anyone except Nadia. Two years ago, after the Business Model Canvas session showed her unit economics that didn’t work, Maya sat at her laptop in the kitchen at midnight and drafted an email to subscribers. “Dear subscribers, we’ve made the difficult decision to pause operations.” Three sentences. Nadia came in and told her to come to bed.

“I never sent it. But I never deleted it either. It’s been sitting in my drafts ever since.”

Dave looks at her. “You don’t farm for the good years. You farm so the bad ones don’t kill you.”

They sit with that. The market fills up around them. A woman with two kids stops at Dave’s stall and buys a pumpkin. Dave waves her off when she tries to pay for a second one. “Take it. They’ll go to waste otherwise.”

Maya watches him. This is what Freshly will never have.

What a feature factory looks like from the inside

The uncomfortable truth about feature factories is that they feel productive. The team is busy. The sprints are full. Features ship. Demos are satisfying. The board sees progress. Subscribers get things they asked for.

The problem is invisible until you measure it. And measuring it requires asking a question that nobody in a feature factory wants to ask: Did any of this matter?

Most teams don’t ask because the answer might be no. And if the answer is no, then the last three months of hard work, the late nights, the sprint commitments, the careful code reviews, were pointed in the wrong direction. That’s not a comfortable thing to confront.

But the alternative is worse. The alternative is continuing to build, continuing to ship, continuing to celebrate output, and slowly drifting from the outcomes that actually determine whether the business survives.

Greenbox drifted for one quarter. They caught it because Charlotte asked a direct question and nobody could answer it. Some teams drift for years.

The headline

The Monday after the market, Sam looks up from her phone at the Perth standup. “Has anyone seen this?”

She reads it out: “Hartland Group acquires Freshly for undisclosed sum.”

The room goes quiet. Freshly, the $12M-funded competitor that has shadowed Greenbox for two years, now belongs to a supermarket group. Hartland’s Richard Ngata, the executive running the acquisition, calls it “a long-term bet on how Australians buy food.”

Maya is the first to speak. “They bought the model. They can’t buy what we built.”

Tom says nothing. He’s thinking about Dave, the call Dave took a year ago when Freshly offered him guaranteed volume. Dave didn’t switch. He stayed.

The standup moves on, because standups do. But the timing sits with Maya all day. Her squads spent a quarter shipping twelve features nobody can connect to an outcome, and the applause had barely faded when a supermarket group bought their closest competitor. Drift is survivable when nobody is chasing you.

The draft

That evening, Maya is at her desk in the Perth office. Everyone else has gone home. The photo of her parents’ farm catches the last light, the one from before it was a subdivision.

She opens her email. Clicks on Drafts.

The unsent email is still there. Two years old. Three sentences she never finished.

Maya reads it once. Then she deletes it. The draft disappears.

She closes her laptop and looks out the window. Somewhere across Perth, subscribers are deciding what to cook for dinner. They won’t have to think about Thursday’s box; it’s already decided. She thinks about the twelve features. Twelve things her team built with care and skill. Two of them connected to a goal. Ten of them built because someone asked and she said yes.

The next chapter, Why, Not Who: The Delivery Morning That Went Wrong, publishes around September 2026.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.