Build vs. buy: a framework for technical decisions
A practical framework for deciding whether to build or buy software: total cost of ownership, core vs. context, switching costs, team capacity, and risk.
By Lance King · · 10 min read
Every team that ships software keeps running into the same question: should we build this ourselves, or pay someone who already has? It shows up for login, billing, email, search, logging, feature flags, and a hundred smaller things. This guide is for engineers, tech leads, and founders who want a repeatable way to answer it instead of rearguing it every quarter. The short version: build the things that make you different, buy the things that just have to work, and be honest about what “build” costs after launch day.
The short answer
Buy if the capability is the same for you as for everyone else (login, sending email, collecting payments), and a mature market of vendors already sells it.
Build if the capability is the reason customers choose you, and an off-the-shelf version would make you look like your competitors.
Buy first, revisit later if you are early, the team is small, and you can keep the vendor behind a thin interface in your own code.
Build (or self-host open source) if a hard constraint rules out vendors: data residency, regulation, unit economics at your scale, or a feature nobody sells.
Either way, write it down in a technical decision record so you know when and why to look again.
What actually matters
Six criteria decide most build-vs-buy calls. You rarely need a spreadsheet with weighted scores. You need to answer each question honestly and notice which way most of the answers point.
1. Total cost of ownership
Total cost of ownership (TCO) is everything you pay over the life of a choice, not just the first invoice or the first sprint. For a purchased tool that means subscription fees, usage overages, integration work, and the time someone spends administering it. For something you build, it means the initial engineering, plus hosting, on-call, security patches, upgrades, documentation, and the slow drift of a system nobody wants to touch.
The build side is where estimates go wrong. Dan McKinley put it plainly in “Choose Boring Technology”: “It is basically always the case that the long-term costs of keeping a system working reliably vastly exceed any inconveniences you encounter while building it.” Google’s SRE book gives a name to much of that ongoing cost: toil, meaning work that is manual, repetitive, automatable, and “devoid of enduring value.” A homegrown system that needs regular hand-holding produces toil every week, and toil tends to grow as your traffic does.
A useful habit: when you estimate a build, estimate year two, not month one. Who is on call for it? Who upgrades its dependencies? What happens when the person who wrote it leaves?
2. Core vs. context
This is the most important question, and the one teams most often skip. Is this capability part of why customers pick you, or is it something every company needs and nobody notices unless it breaks?
Martin Fowler calls this the utility vs. strategic split. Utility functions, like payroll, should be bought, and the business should adapt to the package rather than customize it heavily. Strategic functions are “a crucial part of what makes you better than the competition,” and buying the same software as your competitors “would cripple your ability to differentiate.” Joel Spolsky made the build side of the same argument back in 2001: “If it’s a core business function — do it yourself, no matter what.”
The trap runs both ways. Teams build utilities because building is fun, and they buy strategic pieces because buying feels safe. Fowler also warns about a third mistake: buying a utility and then spending a fortune customizing it, which can cost as much as building it.
3. Time to value
How long until this thing is helping customers? A vendor can often get you to production in days. A build takes weeks or months, and those weeks come out of the work only you can do. McKinley’s “innovation tokens” idea fits here: a company has only a few chances to do something unusual, so spend them on the problem your business exists to solve, not on infrastructure.
Time to value matters most when you are early and still figuring out what customers want. It matters less when the capability is stable and you already know exactly what you need.
4. Switching costs
How hard is it to change your mind later? Jeff Bezos’s 2015 shareholder letter split decisions into “one-way doors,” which are “consequential and irreversible or nearly irreversible,” and two-way doors, which “can and should be made quickly.” Most build-vs-buy decisions are two-way doors if you design for it, and one-way doors if you don’t.
Ask what would stay behind if you left: your data, your users’ password hashes, your customers’ saved cards, years of logs, a codebase full of one vendor’s SDK calls. Open standards help a lot. OpenTelemetry describes itself as vendor- and tool-agnostic, with “no vendor lock-in” for the telemetry you generate. OpenFeature offers a vendor-neutral API for feature flags, so you can switch flag providers without rewriting every call site. Where a standard like this exists, using it shrinks switching costs whether you build or buy.
5. Team capacity and skills
Do you have people who can build this well, and keep it running? A team of four can build an authentication system. The question is whether it should be the team responsible for patching it at 2 a.m. when a vulnerability is disclosed. Count the people, and count their attention. Every system you own is one more thing someone has to understand.
6. Risk
Both paths carry risk, just different kinds. Buying exposes you to vendor risk: price increases, outages you can’t fix, a product being discontinued or acquired, terms changing. Building exposes you to execution risk: it ships late, it has security holes, it works until the one person who understands it leaves.
For some categories, getting it wrong has outsized consequences. Payments, authentication, and anything touching personal data are areas where a specialist vendor has usually seen more attacks, edge cases, and regulators than your team ever will.
A decision table
Instead of comparing vendors, compare the signals. Read across each row and see which column describes your situation.
| Criterion | Points toward buy | Points toward build |
|---|---|---|
| Total cost of ownership | Vendor fees are small next to an engineer’s time to build and maintain | At your volume, usage pricing costs more than a team to run it |
| Core vs. context | Every company needs it; customers never notice it | Customers choose you because of how it works |
| Time to value | You need it this month and the requirements are standard | You can wait, and the requirements are unusual |
| Switching costs | Open standards or a thin wrapper keep you portable | The vendor would hold data or workflows you can’t get back |
| Team capacity | No one on the team wants to own it long term | You have people with the skills and the time to run it |
| Risk | Getting it wrong is costly and specialists have seen more edge cases | A vendor’s outage or price change would threaten the business |
If most of your answers land in one column, you have your answer. If they split evenly, buy now, keep the vendor behind your own interface, and set a date to revisit.
Worked examples by category
These are patterns, not verdicts. Your constraints can flip any of them.
Authentication. For most products, login is context: users expect it to work and never think about it. Doing it well means password hashing, multi-factor auth, social login, account recovery, session handling, and keeping up with security advisories. That’s a lot of risk for no differentiation, so most teams should buy or use a well-maintained open-source option. The switching cost to watch is your user records and password hashes, so check whether you can export them. See choosing an auth provider.
Payments. Moving money touches card networks, fraud, refunds, disputes, and sales tax in many jurisdictions. Almost nobody should build a payment processor. The more interesting decision is how much of the surrounding work to own: tax, invoicing, and compliance can be bought from a merchant of record, a company that sells on your behalf and takes on those obligations. See SaaS payments and merchants of record.
Email. Sending a password-reset email looks trivial. Getting it into the inbox reliably is not: it involves sender authentication, IP reputation, bounce handling, and complaint feedback. Buying a transactional email service is the default. Building your own sending infrastructure makes sense only at very large volume or with unusual constraints. See transactional email providers.
Search. This one genuinely splits. Basic search over a modest dataset can often run in the database you already have. A hosted search service gets you typo tolerance and relevance tuning quickly. But if search quality is your product, as with a marketplace or a documentation site, the ranking logic is core, and you’ll want control over it even if you buy the engine underneath.
Observability. Metrics, logs, and traces are context for nearly everyone, but the bill can grow fast with traffic. A common middle path: instrument with OpenTelemetry so your code is not tied to a vendor, then choose a hosted or self-hosted backend based on cost. See observability on a budget.
Feature flags. A simple on/off flag stored in a config file is easy to build. Percentage rollouts, targeting rules, audit logs, and a UI for non-engineers are much more work. Start simple if your needs are simple, and buy when the requirements grow. Using the OpenFeature API from the start keeps that move cheap. See feature flag tools.
Which one for you
Solo founder or tiny team. Buy almost everything that isn’t your product. Your scarcest resource is time, and every system you build is one you have to run alone. Prefer vendors with free or low starting tiers, clear export options, and standard protocols.
Growing startup. This is where costs start to bite. Usage-based pricing that was trivial at launch can become a real line item. Revisit your biggest bills once a year. Some will be worth replacing with open source or an in-house build; most won’t, once you count the people needed to run the replacement.
Regulated company. Constraints come first. Data residency, audit requirements, and contractual obligations can rule out vendors before cost even enters the picture. Buy from vendors who can show the compliance work you need, and build or self-host where none can.
Company at scale. At very high volume, unit economics can favor building, and you may have dedicated platform teams to run it. Even then, apply the core vs. context test. Being able to build something is not a reason to.
Mistakes to avoid
- Comparing a vendor’s price to a build’s first sprint. Compare multi-year cost on both sides, including on-call and maintenance.
- Building because it’s interesting. Engineers enjoy hard problems. That’s a good reason to work on your core product and a bad reason to write your own job queue.
- Buying and then customizing endlessly. If you’re bending a product into something it isn’t, you’re building anyway, with less control.
- Spreading vendor SDK calls across your whole codebase. Wrap the vendor behind your own small interface. It costs little up front and makes a later switch far cheaper.
- Ignoring data exit. Before you sign, find out how you’d get your data out, in what format, and whether anything (like password hashes) can’t be exported.
- Never revisiting. A decision that was right at ten customers can be wrong at ten thousand.
How to revisit the decision later
Build-vs-buy is not permanent. Plan to revisit it when something changes, not on a whim. Good triggers:
- The vendor’s bill crosses a threshold you set in advance.
- The vendor changes pricing, terms, or ownership.
- You hit a feature limit you can’t work around.
- The capability moves from context to core, because your product strategy changed.
- The in-house system is generating more toil than the team can absorb.
- A mature open-source or standard option now exists that didn’t before.
Write those triggers into the original decision record. When one fires, open a new record that supersedes the old one rather than editing history. The technical decision records guide shows how.
Wardley mapping, a strategy technique created by Simon Wardley, offers a useful lens here. It places each component on an evolution axis from genesis (novel and uncertain) through custom-built and product to commodity (standard, widely available, and expected to just work). Things drift right over time. Something you had to build five years ago may be a commodity you can rent today, and that’s a signal to reconsider.
A quick checklist
- Name the capability in one sentence. Is it core to why customers choose you, or context?
- Estimate year-two cost for both options, including on-call, upgrades, and admin time.
- List hard constraints: regulation, data residency, volume, unique requirements.
- Check data export: what comes out, in what format, and what stays behind.
- Prefer open standards (OpenTelemetry, OpenFeature) where they exist.
- Put any vendor behind a thin interface in your own code.
- Decide who owns it, by name, for the next year.
- Write the decision down, with the triggers that would make you revisit it.
Sources
- Choose Boring Technology — Dan McKinley
- Utility Vs Strategic Dichotomy — Martin Fowler
- In Defense of Not-Invented-Here Syndrome — Joel Spolsky, Joel on Software
- Eliminating Toil — Google Site Reliability Engineering book
- Amazon.com 2015 letter to shareholders (SEC filing)
- Wardley map — Wikipedia
- Landscape: evolution stages — Learn Wardley Mapping
- What is OpenTelemetry? — OpenTelemetry docs
- OpenFeature — project home