Refining the website surface

Refining the website surface

A side quest to let any visitor choose the website’s UI, moving from 7 looks to 93 with 6% more code. Making the website richer on top but simpler underneath.

TL;DR

My last post, “Building the surfaces”, concluded by saying that I was aware of how the website was set up. This post talks about retiring the green skin and letting visitors choose how they view the website.

  • upgrading from 7 looks to 93 looks that a visitor can choose
  • 7 colours, 3 type eras, 12 typefaces from 7 font families, plus 2 fonts for Hindi
  • 1 panel that holds all of it, where pointing at an option previews the whole page and a click keeps it
  • bringing usage of limited admin settings from 11 to 9
  • 6% more website code, with 1,646 old lines deleted along the way
  • 25,748 lines added in total, about 1 in 10 of them the website; the rest were checks and records
  • 45 numbered decisions, 8 build steps in a stable, flexible order
  • done in 3 days - 2 days of designing structure + planning and 1 day of coding

How to read this article

Same as the first post:

  • ‘Tech · ’ pockets hold numbers or short floods of counts for readers who want the details; skim them if you like
  • words with a dashed underline or footnote numbers explain themselves when you hover or tap
  • to jump around, use the floating table of contents at the top on a computer, or its icon in the apps

Before the simplification

In “Building the surfaces” post, I wrote that I was “aware about the rough edges and what I need to simplify and change, not just in the product but also in my process.”

The launch version of the website had four skins, each in two tones, and a simple skin menu that previewed a skin across the whole page when you hovered over it. Some of the points below worked and were actual decisions, but they had noticeable problems:

  • The names led to confusion. Visitors saw Lawn, Paper, Canvas and Glass. The code said classic, solid, pastel and translucent. But “Classic White” was internally solid, while classic was the green one. The AI got confused more than once. A real bug once got chased in the wrong skin because of it.
  • Paper only looked like Paper on Apple devices. Its headings used a font only Apple devices ship while Windows and Android substituted their own.
  • Hindi had no font of its own. It fell back to whatever each device had, so the same sentence looked different on an iPhone, a Windows laptop and an Android phone.
  • The theme asked for font weights in ten different values but the browser drew four. Most of my font weights, it turns out, were suggestions.
  • The homepage wore an accidental colour. It carries strips from several sections, and whichever section won a tie in the stylesheet coloured its frame. This was a frustrating point because it was clearly a structural problem whereas others were easier to fix.

So, I decided to retire Lawn and put Curtain in its place, with the same character, but instead of one fixed green, I would let the visitor pick from a curated set of colours.

I generally start by asking and planning, not by building. By the time the first design session was recorded, three of the early premises that had been made by Claude on medium-effort preliminary code read turned out to need refinement, because they had been made on assumptions. But obviously, things need to be built on real measurements, not assumptions.

There were further complications behind seemingly easy things:

  • Lawn was not “just a green”. It owned a whole-page colour wash made of five soft patches, and effects like a glassy card hover.
  • The wash followed a formula: lightness between 94% and 98%, colour strength in a narrow band, only the hue moving. One seed colour could grow into a whole scheme, which is exactly what a visitor-chosen colour needs.
  • Colouring the site by section cost wildly different amounts in each skin: 0 values in Lawn, 1 in Paper, 4 in Glass and 11 in Canvas.

What visitors get

Simply put, any visitor can now decide how my website looks.

Seven pale colours, named like moods

  • Every colour is pale on purpose. The names are feelings (Plain, Still, Slow, Easy, Cool, Soft, Warm) rather than materials or colours.
  • Visitors choose from a curated set of colors in the panel. Internally, I added guards in the code that ends up refusing to ship two colours that would draw the same swatch.
  • All seven pass the accessibility bar for white text on their colour. Warm is the closest, at 4.58 to 1 against a bar of 4.5.

Typeface by era

  • Instead of the site’s or system’s font, visitors will now have three eras: Legacy, Classic and Modern. The same era can mean a different face in a different skin.
  • Every face is matched to the same letter height, so switching eras never jolts the page’s rhythm.
  • Hindi gets two matching Devanagari fonts, one serif and one sans.
  • Every font now ships with the site itself, so a page looks the same on Apple, Windows and Android.

One panel for all of it

Skin, colour and type now live in one screen-size-aware card that opens from the skin button. The visitor can preview the entire page by hovering/tapping over any option. I kept the light/dark toggle as its own button to save an extra tap/click.

Tech · Before and after, for a visitor

What Before After
Looks a visitor can choose 7 93
Skins Lawn, Paper, Canvas, Glass Curtain, Paper, Canvas, Glass
Typefaces 1 family 12 faces from 7 families, plus 2 for Hindi
Controls for the look a skin menu and a tone button one panel and a tone button
Sections that keep their colour in every skin 14 14

Deciding before building

Much of this project was taking a series of reviews, deciding on choices and system design and planning them out. So, none of my decisions repeatedly touched the website’s code. The method was the same for every decision. I decided after I had the AI render the real page, with the real stylesheet and the real content, including at phone width. The decision went into one written specification that the build was not allowed to contradict.

Tech · What deciding looked like

  • 12 recorded design sessions and 1 review session, with 0 lines of website code changed until the build
  • 12 pages made just for my eyes, about 12 MB of rendered choices
  • 21 colour families rendered on real pages; I shortlisted 7
  • 10 two-colour designs for the colour chips, studied side by side before each chip shrank to one colour
  • 19 font families trialled, in 58 font files
  • every face decided one by one, across a grid of 4 skins and 3 eras
  • 4 computer designs and 2 phone styles for one panel, drawn in place on real pages and checked 117 ways
  • 306 rendered views and 28 assertions at phone width, then 4,298 of 4,298 readings matched against the live site
  • 1 specification, audited against itself before anything was built
  • 165 shared values, 43 data attributes and 21 states mapped, each with every place that writes or reads it
  • 1,094 places where the four skin names lived in the code; an early estimate had counted 15, for one of the names

I can see that look on you again. In plain English: before the build changed a single line, the old website had been measured and mapped in detail.

The last decision: how to build

The last set of decisions were the architecture and build review. Essentially, I needed to be sure that the theme could take all of this as one system, rather than as a pile of special cases. I made the AI map the whole theme first to confirm the five parts would land as one system, with no new special cases, in a fixed order of steps, each proving itself before the next could start.


Building it in eight steps

The build followed the review’s order: rename → frame → defaults → type → colour → panel → zip → go-live. Each step landed with its proofs before the next began, and the build wrote 44 numbered decisions into the decisions log as it went. The full build ran as four back-to-back sessions.

Tech · The build, by session

Session What landed How it was proven
Build 1 The rename, the homepage frame, defaults into code, type by era The rename was reversed and compared with the original build: byte-identical. Old saved skin names fell back to the default 7 times out of 7. 13 of 13 checks for the defaults. Font weights folded into three names. 9 of 9 end-to-end tests.
Build 2 Colour, and the engine that turns seven colours into code 96 of 96 rendered views. 12 of 12 end-to-end tests. 35 new contrast checks.
Build 3 The one panel, then the two changes I asked for after trying it 106 checks, 0 failed. 28 of 28 end-to-end tests.
Build 4 The zip, my eye checks, go-live The zip rebuilt byte-identical to the repository, and the live stylesheet byte-identical after upload.

Every number in that table is a check that came back yes.

Test to build

The panel’s test harness was written first and run against the theme before the panel existed. It failed, as it should: “5 passed, 3 failed”. Then, the panel was built until everything passed. Two of its checks are built to fail forever, as teeth, to prove the harness still bites.

The exam had one especially nasty question. A panel that previews the whole page can shift under your own pointer: the preview changes fonts and sizes, the panel’s contents move, and suddenly the visitor is seeing the page jump around. So I made the AI test the panel with 42 hover events, in each of the 4 skins, on a computer screen and a phone screen, and nothing was allowed to move more than half a pixel.

In the final build session, whose whole job was to prove the final state, the dependency map reported itself stale. My two changes had moved some code, and the map hadn’t been re-run; so thirteen line numbers were out of date. It was a good sign of things working as they should, and in any case, it was a small fix of the satisfying kind. But I think there are better ways to do this.


Code is now simpler, not smaller

The website’s code grew from 14,744 lines to 15,669, about 6%. But simpler in this project wasn’t meant to mean fewer lines. It meant every value is flexible but also lives in one place whenever it needs to, so a change is one edit instead of many, and things didn’t silently disagree with themselves.

When the Curtain work merged into the main line of the repository, it added 25,748 lines and removed 1,975, across 208 files:

  • the website itself (styles, scripts, templates, settings): 2,552 added, 1,646 removed
  • checks and generators: 12,295 added
  • records (the specification, the changelog, the review, the dependency map, the pages made for me): 10,061 added

Tech · Simpler, and heavier on purpose

What Before After
Theme settings in Ghost admin 11 9
Names per skin 2 1
Font-weight declarations 106 3 named weights
Paper’s page colour typed in 6 places 1 shared value
The seven colours did not exist generated from one 29-line file
Font files in the theme (heavier) 8, mostly system fonts 16
What a first visit downloads (heavier) 5 files, about 83 KB, from two places 2 files, about 122 KB, from the site itself
Last font arrives on a simulated really slow phone about 2.4 s about 3.6 s

Working with the AI

This time, I built the website with Claude inside VS Code. I didn’t use Google Antigravity because it would have mostly helped with design variations or checking whether the system design change Claude proposed was good. But Gemini would have consumed a lot of tokens I didn’t want to spend. Also, Gemini wasn’t needed for testing either because I thought that the self-testing process and me being the second pair of eyes were sufficient checks. In any case, Claude is pretty good and following a controlled process to make these changes was also important. So, the 64 commits during this refinement were done through a 60/40 split between Opus and Fable.

Tech · The habits, and what holds them

Habit What holds it Size at go-live
One home for each kind of fact a decisions log that outranks everyone’s memory, the AI’s included 518 lines
a specification the build may not contradict 2,290 lines
one list of what is still open, which must never disagree with reality 189 lines
GitHub issues for discrete bugs and tasks 23 over the project: 12 closed, 11 open by choice
Nothing ships on trust guards, generators and test harnesses 28 script files, 6,281 lines (14 files and 1,986 lines before Curtain)
Every session lands a changelog entry written before each session closes 2,190 lines
The AI remembers through notes a folder of notes that survives between sessions 90 notes today, 45 of them standing rules about working with me
Everything together the current documentation 8,294 lines in 20 files (but I do think I can reduce this severely)

A stranger, or the same AI with no memory, can resume the project from these records alone.

Some new preferences/rules were added

In the last post, I wrote that some of the rules were from learnings and were written after something went wrong. The Curtain project wrote or sharpened 20 of them. The changelog also carries 17 incidents: statements the AI made, then measured, then withdrew. Below is a selection of the patterns:

What happened The pattern it left
I asked what it would take to edit the colours in one place. The AI rebuilt the preview. A question gets an answer, not a build.
I asked whether the admin default settings “can be saved now”. The AI read it as “change the live settings today”, then, once corrected, as “keep them in admin”. An answer at design stage is a recorded decision, never a go-ahead.
The AI initially told me only 2 of the actual 68 pages print Hindi. Count, don’t estimate.
It offered the homepage a “Neutral” option. The preview showed terracotta, warm brown and a peach glow. Render an option before naming it. A fallback is a colour too.
It drew the panel mock-up beside the page, 312 pixels wide. The sidebar it had to live in is 240. Draw every control in place, at real size.
It described a visual difference in a paragraph. Comparisons come as numbered clicks.
I asked what one of its sentences meant. It answered something else and attached a decision to it. Explain first. Ask later.
At go-live, it walked me through the eye checks one question at a time. Batch the checks.

Some general thoughts at this point

What changed most this time was the process, and I designed it because I thought it not only applies to situations when you don’t know the exact mechanics but want to ensure decent-enough quality but it is also extensible to general work that builds on something already live. Some basic working patterns that I like to follow:

  • Decide on the real thing, rendered, never on a description of it.
  • Give every kind of fact one home.
  • Try to write the exam before the work.
  • Let the AI be fast, but make the checks strict.
  • Name what you gave up.
  • Park what can wait, with reasons, somewhere other than your head.

The AI was fast, tireless and, occasionally, confidently wrong in things that would have been silently horrible. But the records being built caught most of that. I caught some of it myself. Neither of us would have managed it alone, and I think that is less a statement about AI than about work in general.

What’s next? The last post ended with me saying that, now that the surfaces were relatively stable, I could focus on sharing my content. I then went and refined a whole bunch of one surface. This time I mean it. Mostly. Building systems forces you to learn.