Skip to main content
Web Development

Content Modelling: Why Page Structure Decides Site Flexibility

a notebook open on a desk with a pen, a phone and printed sheets beside it

The migration that proved the model was wrong

A useful way to understand content modelling is to watch what happens when it was skipped. Take a publishing site that has grown for six years: articles, product pages, campaign landing pages, help articles, staff profiles, and a set of “what we do” pages copied from one another in 2019 because nobody had time to build a template for them. Each is a page, each has its own hand-typed navigation block, and each has its own idea of what a related item is.

printed documents, a clipboard and a pen spread across a desk

Now the business asks for something ordinary. One team should maintain a glossary, and every article should link to the relevant entries. The same author profile should appear on a news page, a conference page and an about page. With pages as the unit of content, each of those is a redesign, and each redesign is a fresh chance for the site to diverge from itself.

With content types as the unit, the same requests are configuration. The glossary is a type. The article has a field that references it. The author is a type with a photo, a title and a biography, reused wherever an author appears.

That is the whole argument. Content modelling is deciding what things are rather than what they look like, so that the same decision only has to be made once.

Model the thing, not the page

The mistake that produces the site described above is thinking in templates. A template answers a layout question: what goes in the header, what goes in the sidebar, how wide is the main column. Templates are necessary. They are not the thing you model.

a whiteboard covered in diagrams and notes in a meeting room

The unit you model is a content type — a thing with a name, an identity and a set of fields. An article has a headline, a body, an author, a publication date and a status. A landing page has a headline, a hero, a form and a set of related links. A staff profile has a name, a portrait, a job title and a short biography. Notice that all three of those share a headline and may share a body. That overlap is a signal that the model has a concept it has not named yet.

Two questions help separate the two kinds of thinking. If you could have fifty of this thing, would they all still be the same type? And if you deleted the page, would the content itself disappear, or only its arrangement? Landing pages for a campaign often fail the second test, because the same hero image and the same four links reappear under a different slug next quarter. That content is not a landing page; it is a reusable block set on a page, and treating it as one produces the copy-and-paste site.

Blocks versus fields

There are two ways to give a type flexibility, and the choice has real consequences downstream.

A field-based type has named slots: hero_heading, hero_image, hero_link. It is predictable, easy to validate, easy to query and easy to report on. It is also inflexible, because every new variation needs another field, and a type with sixty nullable fields is worse to work with than one with eight.

A block-based type has an ordered list of blocks, each with its own settings. A text block, an image block, a related-items block, a form block. Editors gain freedom, and you stop adding fields. The costs land elsewhere: validation becomes “is this combination of blocks legal”, rendering has to handle every block type, and anyone trying to build a search index or a report now has a nested structure to understand.

Most mature systems end up with both: structured types use fields, and the genuinely layout-driven parts of a site use blocks. What to avoid is one “flexible page” type used for everything, because that is how a site ends up with unstructured content that nothing can query.

Relationships, and who owns what

A page-per-page site has no relationships, only hyperlinks. That is why its internal linking rots: the links are typed by hand, so they can only change when somebody remembers. Content models let you declare relationships once and have the site render them everywhere.

Every model needs answers to a few structural questions before it needs answers to anything else.

  • One-to-many. An author has many articles. The article holds a reference to the author; the author is not modified to mention twenty articles. This direction matters. Put the list on the many side and you avoid updating a record every time a new item appears.
  • Many-to-many. An article can have several authors, and an author can write many articles. This needs a join, and in most systems the join needs a field of its own, such as the role of each author on that article.
  • Hierarchies. A help centre with nested sections. Trees are easy to build and easy to get wrong, because they make it awkward to show an item in more than one place. If the navigation is genuinely flat, keep it flat.
  • References versus copies. Whether a page links to a category or duplicates it. Copies look convenient and then disagree with the original. Always reference.

One more structural decision is worth naming because it causes trouble later: whether a piece of content can appear in more than one place. A “trusted by” logo strip that appears on the homepage, on twelve landing pages and in the footer is one object with many placements, not four separate images. The moment you model it as copies, somebody updates one and not the others.

Our custom web application work usually surfaces these relationships as a diagram before any template work begins, because the diagram is far cheaper to change than a working content type with stored data behind it. SmartEdge IT Solutions keeps that diagram alongside the schema definitions, so a developer can see why a field exists without having to ask, which is worth more over two years than it looks at the start.

The route problem: URLs outlive structures

Content models and URL structures are separate decisions that people constantly merge, then regret. The temptation is to mirror the model in the path: /blog/2024/technology/our-post, or worse, /content-type-7/item-4021. Both feel tidy at the time.

shelves of books in a library with a desk and a laptop in front of them

Paths have a longer life than taxonomies. Taxonomies get restructured; marketing reorganises; categories merge. If the URL is a direct expression of the category, every restructure is a redirect project, and redirects decay silently because nobody remembers making them.

The safer pattern is a stable, flat, human-readable path with the taxonomy living in metadata rather than in the address. /guides/how-to-choose-a-database is a reasonable path. The “guides” prefix can stay even if the section is renamed, because nothing about it is derived from a content type ID or an internal category name.

Two things genuinely help at search-engine time, and both are structural. Stable URLs, which mostly means never deriving a URL from a mutable field. And a real 404 that logs the requested path internally, because those log lines are the only reliable source of what people were looking for and could not find. A URL is part of the model whether or not you want it to be, so treat it as an interface with consumers you do not control.

Workflow is part of the model

Teams routinely model content beautifully and then ignore the states it can be in. If a page can be a draft, in review, scheduled, published, and archived, then those states have consequences for every template: whether it appears in navigation, whether search engines can see it, what happens to links pointing at it, and who is allowed to see a preview.

a team working through ideas around a table with notes and a laptop

Scheduling is where a surprising number of sites quietly break. A scheduled item must be invisible before its time and visible after it, without ever passing through a half-published state that gets crawled. Storing the publish time separately from the creation time and applying it at read time is safer than relying on a scheduled job to flip a boolean at midnight.

Permissions belong in the same conversation. If a section is edited by the legal team and needs sign-off, the system should be able to say that; where it cannot, teams find a workaround, usually by giving everyone edit rights and trusting them. Deciding who can change what, and what happens to an item in review when somebody needs to fix it urgently, is modelling work even though it appears in the requirements document.

How much structure is enough?

Over-modelling is a genuine risk, particularly on small sites. If somebody builds a reusable “call to action” type with eight optional fields for a site that has three buttons, they have created maintenance surface and gained nothing. The cost of a well-modelled system is paid on every change; the benefit is paid on the fifth, tenth or fiftieth change. Teams that expect to change rarely should model lightly.

The practical test is whether you can name a concrete change that the structure makes easier. “We want the same author profile on three different page types” is a real reason for an author type. “We might one day want to localise the breadcrumb” is not, and building for it adds a translation dimension to every type you create.

Under-modelling is equally expensive and harder to reverse, because content already stored against a rigid structure cannot be moved without a migration. Once forty pages exist under the old shape, changing the model is a data project rather than a code change.

Both risks point at the same sequencing: model the content you actually have, with the variations that actually exist, in a system that makes the next change cheaper than it would otherwise be. Not the content you might need. Where a business is running a conventional CMS rather than something custom, this is largely a configuration exercise inside a platform such as WordPress or a comparable system; where the content model needs to be unusual, it pushes towards a bespoke build or a headless front end, which is a different conversation about how the content and the rest of the stack connect.

What the modelling session actually looks like

It is boring, and it works. SmartEdge IT Solutions has run this session for sites with three page types and for sites with sixty, and the amount of argument comes out roughly the same in both cases, because the disagreements are never about technology. Somebody from the business brings real examples — actual exported pages, actual PDFs, actual photographs. You list every distinct kind of thing you can find and agree on a name for each. For each, you write down the fields, mark which are required, and mark which are repeatable. Then you draw the relationships between them, and somebody other than the person who drew it checks that the drawing matches what was actually said.

a laptop open on a desk beside a notebook, a phone and a cup of coffee

Arguments happen here, which is the point. The disagreement about whether a case study is a separate type or a variant of an article is a disagreement about who maintains it and what fields it needs. Settling it on paper, with the editor in the room, is the entire purpose of the exercise. Deciding it later means discovering the answer through a support ticket.

Some field types are expensive to change once content exists under the old rules.

  • Rich text. Which elements are permitted, whether heading depth is constrained, whether tables survive a paste. Every platform answers differently, and agreeing that headings are h2 and h3 only is dull work that saves a great deal later.
  • Dates and times. Whether a bare date means midnight local time or noon UTC determines whether a scheduled item appears on the right day for readers outside the business’s own timezone.
  • References to other content. Whether deleting an item deletes its dependents, blocks the delete, or quietly leaves a dangling link. Decide which, then write it down.
  • Anything free text that will be filtered. A field that humans read and a field that machines query need different treatment, and merging them early is awkward to undo.

The honest test for a finished model

A model is finished when somebody who was not in the modelling session can add a new page of an existing type without asking a developer which fields are required. Until that is true, the model is documentation in the form of a database, and the business is paying engineers to make content decisions.

a planner, a notebook and a pen laid out on a desk

Beyond that point, the work shifts to migration, rendering, and how the templates handle combinations the model allows but nobody has used yet. That last category is where content strategy actually happens, because somebody has to decide what a page with no author and no image is supposed to be. Answer it before launch and it is a template fallback; answer it afterwards and it arrives as a bug report from the editor who created it by accident.

Editorial profile

Olivia Bennett Search and Answer Visibility Editor

Olivia Bennett writes for SmartEdge IT Solutions about how people find a service, from technical SEO and structured content through to answer engines and generative search. She reads search console reports for a living and treats an unreachable page as an unfinished one.

Also 2 articles in the Insights archive.

← Back to Blog