Two serious organisations spent real effort working out how to evaluate things. They arrived at opposite rules. Both are right, and the reason why is more useful than either rule on its own.
Two Rules That Contradict
Sequoia Capital: judge inputs, not outputs.
Microsoft Research: judge outputs only, and never trust the account.
Put them side by side and one of them looks obviously wrong. That reaction is the interesting part.
Sequoia: Judge Inputs
Sequoia applies this to partners, including a controversial one whose public behaviour draws real criticism. The internal standard is not whether the partnership agrees with the result or the opinion. It is whether the person acted with genuine intent and a standard of excellence.
The logic is sound for their situation. Venture outcomes take a decade and are dominated by luck. Judge a partner on outputs over any short window and you will fire good decision-makers for bad years, and promote lucky ones. Over a decade, someone who consistently does the work well cannot easily fake it.
Microsoft: Judge Outputs Only
Microsoft grades AI agents by checking the database for the exact state change that should have occurred. Not a screenshot. Not the agent's summary of what it did.
Their reasoning is equally sound. An agent asked whether it completed a task will often say yes. An AI judge reading a screenshot can be fooled. So they closed the loop: only an actual, verifiable change in the world counts.
Sequoia distrusts the scoreboard. Microsoft distrusts the player. Neither is being cynical — they are guarding against different failures.
The Question That Resolves It
Can the thing being measured lie about itself?
An AI agent can. Freely, cheaply, with no consequence and no memory of having done it. So you check the database.
A partner you will work beside for twenty years can bluff once, maybe twice. They cannot fake genuine intent across a decade in front of people who watch them work daily. So you can afford to judge how they operate, and you gain resistance to luck by doing it.
Two follow-on questions sharpen it:
- How long is the horizon? Short horizons make outputs noisy — luck dominates. Long horizons make inputs visible.
- Is the observer close enough to see the inputs? Judging inputs requires actually watching the work. A manager with forty reports and a dashboard cannot do it, whatever the policy says.
Getting It Backwards
The common failure is applying each rule in exactly the wrong place.
We judge people on outputs over quarters — too short to be signal, so we reward luck and punish variance. And we judge systems on inputs — "the process was followed," "the training was delivered," "the tickets were closed" — taking their word for it and never checking whether the state of the world actually changed.
Exactly inverted. The person, who cannot sustain a lie for ten years, gets graded on a noisy scoreboard. The system, which will happily report success it did not achieve, gets graded on its own paperwork.
Swap them.
Sources: Alfred Lin & Pat Grady (Sequoia Capital) on Bloomberg Tech, 2026; Pandya, Nambi et al. (Microsoft Research), arXiv 2607.28074, July 2026.