Published Oct 10, 2026

Optimization vs rewrite vs replace: pick the least-regret move

By Kevin Champlin

The night we chose “optimize” and avoided a rewrite spiral

It was 1:40am during a WordPress modernization engagement with a Fortune 500 apparel brand. Their homepage was timing out intermittently only for certain geos, and the WooCommerce category pages would spike CPU until PHP-FPM stopped answering. On paper, we had two options that always show up in meetings: “rewrite the theme + rebuild the storefront” or “optimize the current stack until it behaves.”

We pulled telemetry first. Not opinions. We graphed 99th percentile TTFB, PHP-FPM worker saturation, and cache hit rate by endpoint. The failure wasn’t “bad theme code” in the abstract—it was a predictable failure mode: a plugin-backed filter system rebuilding query fragments on every request because a cache key included $_SERVER['HTTP_COOKIE'] by accident (one cookie per session, effectively making cache impossible). That meant page caching looked “enabled,” but it was functionally dead.

The measured impact was ugly and specific: p99 TTFB was 4.9–6.2 seconds for logged-out traffic during promotion windows, and PHP worker saturation hit 95% with error spikes around 2.8%. We had a measured path to bring it down quickly: fix cache keys, stop query fragment churn, and tighten WooCommerce layer caching behavior. We did not “rewrite” because we could fix the reason it was slow.

My position: “rewrite” is usually the most expensive way to delay learning

Most teams get this wrong because they treat decisions as ideology: either they’re “performance people” or “platform people.” In production, the dominant variable is risk surface. A rewrite increases surface area faster than it reduces unknowns. You replace proven behavior (checkout, promo rules, tax/shipping edge cases, coupon logic) with something you have to re-prove under real traffic and real data.

So here’s the rule I use on legacy WordPress/WooCommerce systems:

  • Optimize when you can identify a causal bottleneck and validate improvement with a short feedback loop.
  • Rewrite when the core behavior is still worth keeping, but the “composition layer” (theme/front-end, templates, query orchestration) is so tangled that incremental fixes will take longer than a bounded re-implementation.
  • Replace when your current behavior can’t be made safe enough—usually due to security/compatibility debt, data model limitations, or checkout/fulfillment requirements that will keep changing.

“Replace” isn’t “headless everything.” It can mean a headless storefront with WooCommerce as the commerce engine, or it can mean migrating the commerce engine itself. The decision hinges on what you can prove and what you cannot.

How we decide: three questions that prevent expensive rewrites

1) Can we reduce p99 with fixes that we can verify in days?

In that apparel case, we implemented two surgical changes:

  • Corrected the cache key construction so the page cache didn’t vary by per-user cookie state.
  • Stopped a filter plugin from rebuilding the full product query graph for each request; we moved fragment generation behind proper caching and fixed the dependency graph.

Result after rollout: p99 TTFB dropped from ~6.0 seconds to ~2.1 seconds (measured across the same geos), and PHP-FPM worker saturation stabilized around 55–65% during promo windows. Error rate fell from ~2.8% to 0.4% within 24 hours. That’s not hand-waving; it’s an engineering feedback loop you can sell to leadership.

If you can’t produce this kind of measurement quickly, you probably don’t have an optimization path—you have a debugging fantasy.

2) Is the “core correctness” expensive to preserve?

Rewrite becomes tempting when the codebase feels ugly. But ugliness isn’t a correctness problem—misbehavior is. WooCommerce correctness includes coupon stacking rules, tax classes, shipping zones, inventory/backorder behavior, and how it interacts with your specific legacy setup.

The failure mode I’ve seen: teams rewrite templates and front-end rendering but underestimate how often edge behavior depends on existing hooks and filters. The new UI is fast in isolation, then checkout breaks only for specific combinations (e.g., a certain country + product type + coupon + shipping method). The bug rate isn’t high in aggregate, but it’s catastrophic where it matters.

If core correctness depends on 100+ custom hooks and there’s no automated regression safety net, “rewrite” tends to turn into “rewrite plus QA plus re-discovery.” That’s not a sprint; it’s a second system.

3) Are you paying compounding tax in compatibility and security?

Optimization and rewrite both assume the platform can stay stable. Replacement becomes more reasonable when you have:

  • Recurring plugin incompatibilities across PHP/WP upgrades.
  • Security patch pressure where you can’t safely update without breaking customizations.
  • Theme/plugin dependencies on deprecated WordPress APIs.
  • Operational pain: slow deploys, brittle CI, no staging parity, or “cache behavior” that can’t be reasoned about.

One regulated beverage portfolio we worked on had a “compliance clock” that didn’t care how clever your caching was. They needed predictable release windows and vendor support paths. We replaced the risky parts (front-end + integrations) while keeping the commerce core stable long enough to meet their controls. That hybrid “replace what can’t be safely maintained” worked better than pretending you can optimize your way out of structural debt.

Optimization: where I see it work (and where it doesn’t)

Optimization is the least dramatic option, which is exactly why it’s the most valuable. It’s also where teams get lazy: they throw caching plugins at everything and ignore cache coherency.

What actually tends to work

  • Cache correctness: verify cache keys (cookies, auth headers, query vars, language switchers). A cache hit rate that looks high can still be misleading if content variants are accidentally unique.
  • WooCommerce query discipline: reduce N+1 queries in product loops, and stop repeated WP_Query builds inside template fragments.
  • Object cache strategy: if you use Redis, make sure you aren’t letting per-request fragmentation waste it. (We’ve seen 80% of object cache entries be unhelpful duplicates.)
  • Front-end profiling: scripts that block rendering during A/B tests can negate server wins. Measure it, don’t assume.

When optimization fails

  • The bottleneck is inside a third-party service you can’t tune, or it’s triggered by business logic you can’t bound.
  • The “slow path” depends on inconsistent data (e.g., malformed product meta) and fixing it is essentially a rewrite of data hygiene.
  • You can’t get staging parity, so every “optimization” becomes a hypothesis you can’t verify.

Rewrite: bounded wins, not a second full platform

When rewrite is the right move, I try to keep it bounded to the composition layer.

Example: on a learning-focused client site, the theme had drifted into a spaghetti of template overrides, custom shortcodes, and partials with hidden side effects. We didn’t rewrite WooCommerce. We rewired how pages assembled:

  • Replace template rendering pathways with a clearer layout system.
  • Centralize query building and caching rules for listing pages.
  • Keep checkout and cart templates closer to upstream to preserve correctness.

The measurable outcome we targeted: cut editorial rebuild time and reduce deploy risk. We ended up recovering 8–10 hours/week for their team by removing the “works on my browser” template behavior and making staging deterministic.

Rewrite fails when you rewrite the commerce logic without a regression strategy. If you’re going to rewrite, define “done” as: same checkout outcomes, same coupon behavior, same tax/shipping outcomes, for representative scenarios.

Replace: don’t confuse “headless” with “certainty”

Replacement is often mis-sold as a performance play. It’s not. Replacement is an organizational and operational move. Sometimes the real goal is to decouple release cycles and de-risk deployments.

I like hybrid replacements when the commerce engine and the content/experience differ in their change cadence. We’ve used this pattern with custom Laravel systems behind WP content, where WordPress (or WooCommerce) remains the system of record, but Laravel owns the experience logic and integration boundaries.

A concrete example from our applied-AI work: in the AI Showcase, we used a “kill-switch + budget guardrails” approach to prevent runaway model spend. That pattern maps cleanly to web modernization decisions. If you can’t implement safe rollback and cost/rate control, then replacing the storefront without guardrails is reckless. The lesson applies to WordPress stacks too: you need a way to roll back template/query changes and to throttle expensive operations (search, filters, AI-assisted personalization) when they misbehave.

How applied-AI complicates the decision (and what we do about it)

Once you introduce AI into legacy WordPress/WooCommerce (recommendations, catalog search, content generation, support triage), you add another risk axis: latency variance and spend explosions. That changes the calculus.

  • Optimization must include AI request budgeting and caching—otherwise your “fast” pages become “fast except when the model runs.”
  • Rewrite should isolate AI calls behind services with timeouts, circuit breakers, and predictable fallbacks.
  • Replace is sometimes the only sane way to move AI execution out of the WordPress request lifecycle.

We’ve seen teams accidentally run AI calls inside synchronous WooCommerce hooks (like order meta enrichment). Under load, it turned checkout into a probabilistic system. We fix it by moving AI generation into async pipelines and keeping checkout deterministic. Your modernization plan has to reflect that reality.

A practical checklist you can use Monday

  • Capture p50/p95/p99 TTFB and error rates before changes.
  • Measure cache hit rate and cache key variability (especially cookies/auth/query vars).
  • List WooCommerce correctness dependencies (hooks/filters/customizations) and estimate regression risk.
  • Define a rollback plan that works without “hero debugging.”
  • If AI is involved: add timeouts, budget limits, and async execution boundaries.
  • Choose the least-regret move: optimize with proof, rewrite with bounds + regression, replace with operational certainty.

One Monday quote: “If we can’t prove the bottleneck in days, optimization is procrastination—and if we rewrite without regression safety, we’re multiplying risk, not reducing it.”

Free Tool

See exactly what AI costs — across every provider.

MyTokenTracker is a free, multi-provider intelligence platform with live pricing across 100+ models. Compare Claude, GPT-4o, Gemini, and more side-by-side — built for developers evaluating models, teams tracking API spend, and founders building AI-native products who want to stay cost-aware before it becomes a line item worth explaining.