Skip to main content
Web Development

What Core Web Vitals Mean for a Site You Own

a laptop and a monitor showing charts and performance reports

Core Web Vitals are described as user experience metrics, which is broadly right and slightly misleading. They are not measurements of whether a page feels good. They are three measurements of specific failures — a main element appearing late, the page moving while you try to read it, and the interface not responding when you touch it — aggregated across real visits. Understanding that they are separate failure detectors, rather than one score, is what makes them actionable. Most disappointing performance work happens because somebody treated the numbers as a single number.

Three metrics, three different owners

Largest Contentful Paint measures when the largest element of the visible page finished rendering. It is standing in for “did anything meaningful appear quickly”, and it is dominated by whatever happens before that element: server response time, the discovery of the resource that carries it, and how long the browser took to decode and paint it.

a smartphone held in one hand with an app open on the screen

Cumulative Layout Shift measures how much visible content moves after it first appears. It is standing in for “was the page stable enough to read and click without it moving under you”. It has nothing to do with how fast anything loads — a page can score well and feel slow.

Interaction to Next Paint measures whether the interface responds promptly after a tap, click or key press, by looking at whether something new is painted afterwards and how long that takes. It is standing in for “is this thing alive when I touch it”.

Each has a different root cause and a different person who can fix it. LCP is usually a back-end or asset-delivery problem. CLS is almost always a layout problem, and frequently involves a font or a reserved slot. INP is usually front-end JavaScript, and it is the hardest of the three to improve without removing functionality. Treating them as one budget means the team optimises whatever is easiest and the headline number moves very little.

Field data is the score; lab data is the diagnosis

Field data comes from real visits on real devices, and it is what the thresholds are applied to. Lab data comes from a simulated run on a machine you control. Only the first is a score; the second is a way of finding out why.

a laptop open on a desk beside a notebook, a phone and a cup of coffee

Lab measurements are consistently better than field measurements for the same page, because the devices, connections and processor contention in the field are worse than anything you would configure deliberately. This gap is not a mistake in the tooling, it is the whole point. The variation between your own test device and a mid-range Android phone on a congested network is the actual problem you are trying to solve, and no amount of lab work will close it.

Field data also has a coverage problem. It requires enough traffic to the page type to be aggregated, so a long-tail template or a low-traffic campaign page may never produce a verdict at all. Low-traffic pages therefore need lab work and careful engineering judgement rather than waiting for data.

The two mistakes people make with these numbers

The first is chasing a lab figure. Running a test on a developer machine, getting a good result, and concluding the site is fine is the single most common way this work becomes wasted. The second is treating the aggregate as a verdict on a specific page. An estate can pass comfortably overall while one template — usually the busiest, or the one nearest a commercial goal — is well outside the threshold and dragging the rest.

Where a site is genuinely small and has very little traffic, the sensible response is to ignore the aggregate entirely and treat the thresholds as engineering targets, instrument the pages that matter yourself, and be honest that you are reasoning from lab work. At SmartEdge IT Solutions we would rather state which of the two we are doing than report a percentage from a tool that has never seen a real visit.

Field data is also historical. You are looking at what happened over a recent window, which means the newest release may not be represented yet. It is entirely possible to fix LCP on a template and see the number not move because the window still contains a month of the old version. This is the usual reason performance work gets abandoned, and the fix is to check the data split by page version and template before concluding anything.

LCP: usually the server first, then the image

When LCP is poor, resist starting with images. Work backwards from the element: how long until the server responded, how long until the resource was requested, how long until it had downloaded, how long until it had decoded and painted. Each of those has a fixable cause, and the order matters because fixing a later stage while leaving an earlier one untouched produces no change.

a laptop open on a desk beside a notebook, a phone and a cup of coffee

Server response time on a dynamic page is the usual first culprit. A request that spends several hundred milliseconds assembling a page from a database with several queries has already spent a large part of its budget before any element can paint. Caching at the edge, or rendering the page ahead of time when the content does not change per request, addresses this directly. It is infrastructure work, which is why DevOps involvement tends to matter more here than frontend optimisation.

The next stage is discovery of the LCP resource. If it is a hero image referenced in markup, the browser finds it almost immediately. If it is injected by JavaScript after a bundle has parsed and run, there is a delay you did not account for. That delay is frequently larger than the image download itself.

Only then do images come into it, and then the fixes are unglamorous and effective: modern formats rather than what the design software exported, sensible dimensions with the correct aspect ratio, responsive sources so a phone does not download a desktop-sized file, and preloading the one element that will be the LCP on that template.

CLS: reserve the space, do not animate it away

Layout shift happens when something arrives after the first paint and takes up space that was not reserved. The canonical causes are unglamorous: images without width and height attributes, a banner or consent bar injected at the top of the page after load, embedded content with unknown dimensions, and — the one that catches most people out — web fonts.

Web fonts cause shift because the fallback font used for the first render has different metrics from the font that eventually loads, so every line of text reflows. The fix is either to size the fallback to match the metrics of the real font, or to use the platform font stack and skip the custom font on pages where it does not earn its place. Both require a design decision about how the loading experience looks, which is why design input belongs in this conversation rather than only development.

The other recurring source is anything injected by a third party. A consent dialog, an ad slot, a chat launcher. The available strategies are to reserve space for the element, to position it so it does not move anything above it, or to load it after the page has settled — and in practice the honest answer is that a slot which changes size unpredictably will keep causing shifts until somebody designs it properly.

One caution about the measurement itself: CLS has a session window, and short sessions with a single interaction are excluded. If your site produces a lot of very brief visits, your score will look better than the experience, and it is worth looking at the shift events directly rather than only at the aggregate.

INP: the honest one

Of the three, INP is the metric most likely to be affected by a decision somebody made deliberately, and the one where we most often end up recommending that a client accept a number rather than chase it.

office buildings and towers photographed against an open sky

The cause is nearly always main-thread work: a large JavaScript bundle parsing and executing, a framework re-rendering more of the tree than the change required, a long-running task from a third-party script, or an event handler that does substantial work synchronously. The standard techniques help. Code splitting so the initial bundle is smaller, avoiding work in event handlers that could be deferred, moving computation off the main thread, and removing libraries that duplicate what the platform already does.

The catch is that most of those changes either remove functionality or add complexity, and neither is free. A team that ships a rich interface and a team that ships a fast interaction are often making the same trade-off differently. This is where we spend time in review rather than in a sprint: which features on this page are actually used, which can be deferred to after the interaction completes, and which would be better as a server-rendered page with a smaller amount of live behaviour.

It is also worth noting that INP is measured across interactions on a page, so a single heavy component — a data table, a map, an editor — can dominate a whole template. Measuring per interaction rather than per page is the only way to find it.

The trade-offs, stated plainly

Every performance improvement takes something. Larger caching means content can be briefly stale after an edit. More aggressive image optimisation means a manual review before it goes anywhere near brand assets. Prefetching means bandwidth spent on pages some visitors never open. Deferring JavaScript means a slower first interaction for the feature you deferred. Code splitting means more requests.

a close view of a desk with a keyboard, a notebook, a pen and a coffee cup

These are not obstacles to be overcome with effort; they are the actual decisions. The useful thing a performance conversation produces is a written position on which trade-offs the business accepts. For most content sites, brief staleness after an edit is a cost nobody minds, which makes aggressive caching easy to justify. For a pricing page or a live dashboard, it is not, and that should be stated before the cache policy is written.

Another point worth making early: a large share of the largest single cost on many sites sits in third-party tags that nobody owns — analytics, chat, heatmaps, personalisation, advertising. Each is individually small and collectively decisive.

Tag governance is not a performance fix

This is a commercial decision, not an engineering one, and it needs whoever signed the contracts involved. Removing a tag can invalidate a contract, break a reporting dependency or remove data somebody is measured on, which is why engineers should not be handed the authority to delete them unilaterally. What engineering can do is make the cost visible: a per-tag inventory with its transfer weight and main-thread time, so the conversation is about numbers rather than about whether the page feels heavy. That inventory is part of monitoring and tag ownership work, not a performance ticket, and it tends to outlive any one optimisation round.

What to check before spending anything

Before optimising, establish the baseline and its limits. Which templates carry the most traffic, and how many are there? Is there field data at all for the pages that matter commercially, or are you inferring from a lab run? Which single change is likely to move the largest share of real visits — because on most sites the answer is one template, and the other twenty can wait.

a laptop open on a desk beside a notebook, a phone and a cup of coffee

Then check the things that are free. Uncompressed assets, a missing cache header, an unoptimised hero image on the busiest landing page, an analytics snippet loaded twice because two teams added it. These are common, they are quick, and they often account for most of the achievable improvement on an estate nobody has looked at in a while. An audit on a site nobody has measured in years usually finds several of these in its first hour.

Finally, decide what you are willing to keep. A site that scores well because it does less is a better site for most businesses, but it is a decision about the product, and it should be made deliberately. SmartEdge IT Solutions produces a ranked list of changes with an estimated effect per template rather than a single score, because that is the only format that helps anybody decide. If you want that list for a site you own, get in touch and we will tell you what it is likely to contain before you commit to anything.

Editorial profile

Emily Carter Web Development Editor

Emily Carter edits SmartEdge IT Solutions articles on website builds, content management systems, storefronts and the maintenance that follows a launch. She is interested in the parts of a web project that decide whether it is still easy to run two years later.

Also 2 articles in the Insights archive.

← Back to Blog