Comparisons
Pairs that get conflated in real conversations and in real pull requests — coupling and cohesion, abstraction and indirection, refactoring and rewriting, debt and mess. Neither column wins; what decides is the requirement. Each record leads with the confusion, because the confusion is the reason the record exists.
Exceptions vs Result types
The debate is usually conducted as a language-culture fight — Go versus Java, Rust versus Python — when the useful distinction is about which failures are *part of the domain*. A declined payment is a business outcome that every caller has an opinion about; a null dereference is a bug that no caller can meaningfully handle. Putting the first in an exception hides it from the type system and from the reader, so the handling ends up wherever someone remembered; putting the second in a return type forces a hundred call sites to propagate something none of them can do anything about, which is how result types become noise and get ignored with an unwrap. The second confusion is a practical one about cost: exceptions are invisible in signatures, which means an interface can grow a new failure mode without any caller being told, and checked exceptions were the attempt to fix that which most languages abandoned because it made interfaces unchangeable. Result types make the failure visible at the cost of ceremony in every frame between the failure and the decision. Most well-designed codebases use both, deliberately, with a written rule about which is which — and the rule, not the choice, is what makes the codebase predictable.
For failures that mean the program's assumptions are broken — bugs, impossible states, unavailable infrastructure — where every intermediate frame would only rethrow and the right handler is far away at a boundary.
For outcomes that are part of the domain: the card was declined, the coupon expired, the name is taken. These are not exceptional, callers must decide about them, and the compiler should insist that they do.
| Dimension | Exceptions — failure unwinds the stack to a handler | Result types — failure is a value in the return type |
|---|---|---|
| Visible in the signature | No, in most languages | Yes — it is the return type |
| Suited to | Faults and infrastructure failures | Expected domain outcomes |
| Cost to intermediate frames | None — they do nothing | Explicit propagation at every level |
| Compiler can insist you handle it | Only with checked exceptions, which most languages dropped | Yes, when the type is not silently discardable |
| Failure mode | Caught too broadly, logged, swallowed | Unwrapped without checking, becoming a worse exception |
| Effect on interface evolution | A new failure can appear without callers knowing | A new variant is a visible, breaking change — which is honest but costly |
| Ergonomics | Excellent for the happy path | Depends heavily on language support for chaining |
| The real decision | Which failures are bugs | Which failures are business outcomes — write the rule down either way |
The same question, five structures
Layered, hexagonal, clean, vertical slice and modular monolith — compared without naming a winner, and with the block that says where the comparison stops being true.
A tidy table implies an equivalence that does not exist. These are not five points on one axis: layered, hexagonal and clean are statements about dependency direction; vertical slice is a statement about directory grouping; modular monolith is a statement about deployment and module visibility. Most real systems combine several. The where this comparison misleads block on every row is the part worth reading, and it is the reason this page names no winner — none of these is mandatory, and a team that adopts one because a diagram was pretty has skipped the only question that decides it.
These are not five points on one axis. Layered, hexagonal and clean are all statements about *dependency direction*; vertical slice is a statement about *directory grouping*; and modular monolith is a statement about *deployment and module visibility*. You can — and most real systems do — combine several of them: a modular monolith whose modules are vertical slices, each with a hexagonal boundary at its edges. Comparing them as alternatives is the single most common way this table is misread.
The word complexity is doing two jobs here. Layered and vertical slice are cheap to *set up* and can be expensive to *live in* once the codebase is large; clean and hexagonal are expensive up front and their cost is roughly flat afterwards. Any comparison made at week one inverts the ranking you would get at year three, and neither reading is dishonest — they are answering different questions.
Locality is a property of whether the boundaries match the change history, not of the style name. A vertical slice cut along the wrong capability lines has terrible locality, and a layered codebase with only one real feature has perfect locality. The only honest way to compare these columns is to open the last thirty merged changes in your own repository and count the directories each one touched.
Every column here is a claim about *fast tests without infrastructure*, and any of the five achieves that as soon as dependencies are injected rather than constructed — which is a separate decision none of these styles owns. What differs is the default test boundary each one nudges you toward, and that matters more than the theoretical maximum: layered nudges toward class-level tests with mocks, vertical slice toward feature-level tests, and the difference shows up in how much your suite has to change during a refactor.
The columns are answering to different pressures: hexagonal responds to *external* volatility, clean to *domain* richness, vertical slice to *feature count*, and modular monolith to *team count*. A system can score high on one pressure and low on the rest, which is why picking a style from a comparison table rather than from your own pressures is how teams end up with four rings around a CRUD application.
Team fit is not a tiebreaker, it is often the deciding factor, and it is the one this table cannot capture. A structurally superior design that the team will not maintain under deadline degrades into the worst version of itself — half-applied clean architecture, with some code respecting the ring rule and some not, is harder to work in than consistent layering. The right question is which of these your team will still be following in eighteen months.
This row compares familiarity, not intrinsic difficulty, and familiarity is a property of the industry at a moment in time rather than of the design. Layered wins here largely because it is what most people have seen, which is an argument for it and also the reason it is over-applied. It is also worth separating cost-to-read from cost-to-contribute-correctly: vertical slice inverts on those two, and the table's single number hides it.
Ceremony is only waste when the feature did not need it, and every column here is right for some features and wrong for others in the same codebase. That is the actual finding of this row: a uniform ceremony level applied to every feature guarantees you are overpaying on the simple ones or underpaying on the complex ones. Allowing different features to carry different amounts of structure is more valuable than choosing which column to standardise on.
shared/ directory that grows back under a new name, or slices that each reimplement infrastructure slightly differently.Every one of these degradations is the style's own strength taken past the point where it repays — which is why none of them can be called wrong, and why §138 forbids teaching any as mandatory. What makes a codebase bad is not the column it started in but the absence of anyone asking whether the structure still matches the changes arriving. The right comparison to make is between your current structure and your last thirty changes, not between two names on a page.