A north star metric is the single number that best captures the value your product delivers to customers, used to align a whole company on one definition of progress. The hard part is not calculating it. It is getting a leadership team to agree on it, because whichever number you pick decides whose work counts as progress and whose work becomes a supporting cost.

The whiteboard had two numbers on it. Total leads and qualified leads. Everyone in the room had already agreed, in principle, that EverQuote needed one number to run the expansion against. We were taking an auto insurance marketplace into home, life, and health. Four verticals, one company, one definition of winning. The agreement held right up until we had to pick which of those two numbers it was.

That gap between agreeing in principle and picking the actual number is where the real work sits. Every article about how to find your north star metric treats the choice as an analysis. Study the product, find the moment of value, write it down. I have never seen that be the hard part. The hard part is the room.

What is a north star metric?

The definition itself is settled and uncontroversial. Sean Ellis, who popularized the framework, set three tests for the number: it should reflect real customer value, teams should be able to move it through their daily work, and it should predict revenue without being revenue itself. Anything failing those tests is a KPI with a better title.

The canonical examples are clean. Airbnb counts nights booked, which captures value flowing to guests and hosts at the same time. Facebook counted daily active users. Both numbers go up only when the product is doing the thing it exists to do. Ellis adds one more constraint that gets ignored more than the others: the metric should not be a ratio. Ratios can improve while the business shrinks, and a company chasing a ratio will eventually find that out the expensive way.

Those criteria are good. They are also exactly why the number gets contested. Any metric that predicts revenue is a metric people will be measured against, and people know it before you finish drawing the box around it.

Every article about how to find your north star metric treats the choice as an analysis. The hard part is the room.

Why is picking a north star metric a negotiation?

Because the number redistributes power. Whichever metric you choose decides which team's work shows up as progress and which team's work turns into a line item that supports someone else's progress. It is a budget conversation wearing an analytics costume, and everyone in the room can see through the costume.

Look at what was actually on that whiteboard. Total lead volume and qualified lead volume sound like cousins. They are not. Total lead volume is a number you can move by buying more traffic. Qualified lead volume only moves when a consumer finds coverage they actually want and a carrier receives a lead that converts. One is an input you can purchase on a Tuesday afternoon. The other has to be earned on both sides of the marketplace.

Follow each choice to its org chart. Pick total lead volume and paid acquisition becomes the engine of the company, with product in a support role feeding it landing pages. Pick qualified lead volume and matching, consumer experience, and carrier integration all become first-class work with first-class budget. Same company, same week, two very different companies implied by two words on a whiteboard.

The team picked qualified. That decision is a large part of why the multi-vertical expansion worked, and why the business went from roughly $200M to roughly $400M over those years. It is also the reason the seven product metrics I keep coming back to all share one property: each of them is expensive to move dishonestly.

📌
If nobody in the room objects to your north star, you have not picked a real one. A metric with no losers is a metric with no teeth.

How do you run a north star metric workshop?

Put every serious candidate on the board and make the room answer the same four questions out loud for each one. What moves this number. Who wins if we pick it. How would someone game it. What do we stop doing. The disagreement lives in the answers, and the answers take about two hours.

Run those four questions against the two candidates from that whiteboard and the shape of the decision comes out fast.

Question asked of each candidateTotal lead volumeQualified lead volume
What moves this numberBuying more trafficBetter matching between consumer and carrier
Who wins if we pick itPaid acquisitionProduct, data science, carrier integration
How someone games itLoosen what counts as a leadTighten qualification until volume stalls
What we stop doingPaying for match quality we cannot seeCounting volume that never becomes coverage

The fourth question does most of the work. What do we stop doing is where people who have been nodding along for forty minutes suddenly discover they have opinions. Everyone will agree to a metric in the abstract. Almost nobody agrees to the specific thing that metric takes away from them, and you want that fight in a room with a whiteboard rather than six months later in a performance review.

The third question matters nearly as much, because a metric you cannot game is usually a metric you cannot move either. Naming the exploit out loud tells you where the guardrails go. Qualified lead volume can be gamed by tightening the qualification bar until the number looks clean and the business quietly starves. Once the room says that sentence out loud, someone owns watching for it.

Run this before anyone builds a dashboard. Once the dashboard exists, the metric has been decided by whoever built it, and the conversation you are having is no longer a decision. It is a review of someone else's decision.

Whatever number survives, it earns its keep in the weekly rhythm rather than in the deck. The number belongs at the top of the weekly product review, where the same people look at the same figure often enough to notice when it stops telling the truth.

What makes a north star metric fail?

It fails when it becomes a target. Charles Goodhart, then advising the Bank of England, wrote the original version in 1975: any observed statistical regularity will collapse once pressure is placed on it for control purposes. The anthropologist Marilyn Strathern later compressed it into the phrasing everyone quotes: when a measure becomes a target, it ceases to be a good measure.

The Bank's own story is the cleanest illustration available. It had noticed stable relationships between certain measures of money supply and inflation. It started targeting those measures to control inflation. Banks and markets adapted to the new rules, and the reliable correlation evaporated. The relationship was real right up until it was load-bearing.

Every north star is a statistical regularity somebody noticed. Nights booked correlates with Airbnb being healthy. Qualified lead volume correlates with a marketplace being healthy. The moment you hang the whole company's compensation on that correlation, people start pressing directly on the number instead of on the thing the number was quietly measuring. John Cutler, who co-wrote Amplitude's North Star Playbook, has a good test for this: if you can move your north star directly, it is probably not a good north star.

⚠️
A north star metric you can move directly is a dial, not a north star. If one team can hit the number without any other team changing behavior, you picked an input and called it an outcome.

The failure mode I see most often has nothing to do with the metric being wrong. It is a company that adopted a good number without ever having the argument. At EditMe we ran into a version of this early, and the whole thing turned on defining activation correctly before we built anything to measure it. The definition was the product of the argument. The dashboard was just where we wrote the answer down.

At EnergySage I watched the same physics from the CPO seat. A two-sided clean energy marketplace has an obvious temptation: count the thing that is easy to count. The number of quotes requested is easy. Whether a homeowner ended up with panels on the roof is hard. The easy number moves on Monday. The hard number is the business.

What the metric is actually for

A north star metric is a coordination device. It exists so that a product manager in one building and a data scientist in another can make the same call on a Tuesday without a meeting. That only works if both of them were represented in the argument that produced it, or at least trust the people who were.

This is why importing another company's north star never takes. Nights booked works at Airbnb because Airbnb fought about it. Copied onto your wall, it is decoration. The number has no memory of your constraints, your org chart, or the thing your team agreed to give up to make it true.

Which points at something uncomfortable about how what people are rewarded for drives everything downstream of it. When a team is not moving together, the instinct is to restate the goal more clearly. Usually the goal was stated fine. The reward structure underneath it was pointing somewhere else, and no amount of restating fixes that.

Companies that skip the argument and adopt the metric get the dashboard without the alignment, then wonder why nobody changed what they did on Monday.

The argument is the deliverable. The number is just where you write the answer down. When people ask me how to pick a north star metric, they are usually asking for the criteria, and the criteria are already published in a dozen places by companies that sell dashboards. What they are missing is two hours, a whiteboard, and the willingness to make the room say out loud what it is prepared to stop doing.

Two words on a whiteboard. Qualified, not total. It took an afternoon to choose and it set the direction of the next two years.