Files
adr-sml/docs/04-seo-spec.md
T
Pouya LajevardiandClaude Opus 5 210bc25a26 feat: build steps 7a-10 — the site is complete and reviewable at 22 pages
Steps 7a through 10 as one authorised run. Nothing deployed (D11).

7a  Lighthouse returns as `lighthouse@13.4.1` + `chrome-launcher`, NOT
    `@lhci/cli`. AGENTS.md §7's advisory attribution was wrong: the carriers
    were @lhci/cli's own `tmp` and @puppeteer/browsers' `extract-zip`, not
    Lighthouse, which audits clean. A deliberate deviation from R11's literal
    trigger, recorded with what it costs. Local gate; CI has no Chrome.

7b  OG card generator (satori + sharp) discharges R15 — 20 typed cards plus
    per-article cards; the portrait stays on / and /about/ by Q40. Insights
    plumbing: ArticleCard, Prose, the index, the article route, articleGraph,
    and /'s section 7. Card copy is constrained structurally because text in a
    JPEG cannot be grepped by check:claims: every headline IS its page's <h1>,
    enforced by `npm run og:proof`.

7c  Five drafted launch articles, draft: true / reviewedByPouya: false. An
    independent compliance audit returned 76 findings and 57 unsourced
    assertions; all blocking and should-fix applied.

8   /contact/, the intake form, and backend/intake/ (undeployed). Plain HTML
    POST to a same-origin /api/intake with a 303 redirect, so the form works
    with zero JavaScript. docs/05 records three deliberate deviations.

9   /fees/ on Q59's ruling — overtime runs from the session cap, and the
    reservation point ships adjacent to the rate. One-page PDF bio discharges
    R16; /bio/ is its source, so the circulated artefact stays inside the
    review apparatus.

10  /legal/privacy/ and /legal/terms/, written to the backend as built. Three
    of the policy's statements are derived and cannot drift.

Also: /about/'s inverse credentials band (approved at step 6); Q59 closed;
R15 and R16 discharged; and a fix to shipped copy — /practice/energy/ asserted
the absence of a regulation the source extract says must not be asserted.

Review: adversarial-reviewer, two rounds (D20/D19). Round 1 returned 16
findings including two blocking — an invisible ghost button on /fees/ at
1.00:1 that Lighthouse scored 100, and a privacy policy that named one data
processor when there are two. All 16 acted on.

Lighthouse, 22 pages, mobile: performance 99-100, accessibility 100,
best practices 100, SEO 100 on every indexable page, CLS 0.000.

AGENTS.md entry (ah) has the detail, including four of my own verification
commands that were wrong and what each of them nearly caused.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Md3GndFqWPzK78xAoebsg5
2026-08-31 10:56:54 -04:00

17 KiB
Raw Blame History

04 — Discoverability

The problem this project exists to fix. AGENTS.md §2 has the measurements.

Those measurements are under review — AGENTS.md Q34. They were taken on 2026-08-25; a re-fetch on 2026-08-26 returned a bundler harness whose real <head> sits JSON-escaped inside a <script> and whose application lives in nine UUID-named files that were not fetched. Some of §2 reproduced exactly (the 2.2 MB single file, the placeholder <title>); some could not be reproduced from the served HTML at all. Cite §2, and cite Q34 with it. Do not put any of these figures in public copy until Q34 closes.


The baseline being replaced

Now [verified 2026-08-25] Target
Content in server HTML SML Company · DISPUTE RESOLUTION · Unpacking... Every word
Indexable pages 1 17 + articles (19 fixed URLs, less the two /legal/* pages, which are noindex and excluded from the sitemap)
<title> SML Company · Dispute Resolution — pre-rebrand placeholder Unique per page
Meta description none Unique per page
<meta viewport> absent Present
Canonical URL none Every page
OG / Twitter tags none Every page
Structured data none Person, ProfessionalService, Article, FAQ, Breadcrumb
robots.txt 403 Served
Sitemap none Generated at build
Favicon none Full set

Astro's static output solves most of this by existing. The rest is the spec below.

Why static output matters more than usual here

Google can sometimes render client-side JavaScript. Bing largely does not. LinkedIn's preview crawler does not. Slack's unfurler does not. And the crawlers behind AI assistants — increasingly how counsel and in-house teams find a neutral — generally do not.

A site that requires three CDN round trips and an in-browser Babel compile before producing a sentence is invisible to all of them. That is the whole argument for D1.


Metadata

Every page passes through one SEO component. A page without it is not finished.

title            5060 chars, unique. Pattern: "<Page> · Pouya Lajevardi".
                 `/` composes its own from the constants — `SITE.name` +
                 `SITE.tagline` — rather than a literal, so the masthead and
                 the title cannot drift. This spec carried the literal with an
                 ampersand after the composition shipped with interpuncts;
                 cite the constants, do not restate them (AGENTS.md §7 rule).
                 ARTICLES ARE THE EXCEPTION: no " · Pouya Lajevardi" suffix.
                 The suffix is 18 chars, so a headline that already reads
                 5060 renders at 6878 — over this ceiling. Measured against
                 the five launch headlines in 03-content-spec.md, the suffix
                 rule fails 5 of 5; without it, 4 of 5 pass. An article's
                 headline IS its <title>; `seoTitle` in the frontmatter
                 overrides it when a headline that reads well is out of range.
                 src/content.config.ts enforces this and names the offending
                 string and its length in the build error.
description      140160 chars, unique, written for a human, not stuffed
canonical        absolute, https, trailing slash
og:title/description/image/url/type/site_name/locale  (en_CA)
twitter:card     summary_large_image
robots           index,follow — except /legal/* which is noindex,follow

OG images: 1200 × 630. Never a screenshot.

RULED 2026-08-27 (AGENTS.md Q40, R15) — TWO kinds of card, not one, and the generator is deferred to build step 7. This spec said "one template" for all nineteen pages. Pouya split it:

"A portrait is the right OG image for / and /about/ — a face is the strongest social preview for a personal brand. It is the wrong one for nineteen pages, where a typed card carrying the page title would do the work.

But do not build the generator now and do not leave 'portrait everywhere' as an untracked interim. Ship it at step 7 alongside Insights, which needs per-article cards anyway — one build, one dependency, one review."

So:

Pages Card
/ and /about/ The portrait crop, src/assets/og-portrait.jpg. Not an interim — the decided answer. Resolved from PORTRAIT_PAGES in src/data/og-cards.ts, not from a per-page prop
Every other page Generated at build by src/pages/og/[...slug].jpg.ts from satori + sharp, in the site's own type and palette: display headline on cream, infinity mark, designation line
Each article Per-article card from the same endpoint — the reason the two jobs were one build

BUILT — step 7b, 2026-08-31. R15 IS DISCHARGED. satori@0.33.4 was chosen over astro-og-canvas@0.13.0 (both 0 vulnerabilities, verified that day): sharp is already a dependency to rasterise satori's SVG, so it adds one library rather than a CanvasKit wasm blob, and it renders with this site's own fonts and tokens rather than approximating them.

Four things about the implementation are load-bearing and are not style choices. Each is recorded because a later reader would otherwise "tidy" it:

  1. Colours are parsed out of src/styles/tokens.css at build time, not copied into the generator. CLAUDE.md requires every colour to come from a token; the alternative was a duplicated hex table, which is the SES-DKIM shape. A missing token throws rather than falling back.
  2. The fonts are @fontsource's static .woff cuts, not public/fonts/. satori parses TTF/OTF/WOFF and not WOFF2, and decompressing the site's own subset variable Geist to TTF throws inside satori's opentype.js fork — Fontsource's subsetting drops the name records the fvar table points at. Same typeface, same upstream version, same weight; build-time only.
  3. Every card's headline is its page's own <h1>, character for character, and npm run og:proof enforces it against the built HTML. This is a compliance mechanism, not a convenience: text baked into a JPEG cannot be grepped by npm run check:claims, which under D20 is the only per-step claims control there is. A card must not carry a claim its page does not already make in auditable HTML. The same check confirms every page's og:image resolves to a file that exists — a 404 preview is invisible from inside the repo.
  4. A page with no card entry is a BUILD ERROR, not a fallback to the portrait. R15's failure mode was never the wrong image; it was the wrong image shipping invisibly and reading as intentional. A silent fallback recreates it exactly.

Structured data

JSON-LD only. Validate against Google's Rich Results Test before cutover.

Type Where Notes
Person /about/, referenced site-wide Emitted: name, url, jobTitle, description, alumniOf (Bond University), knowsLanguage (en, fa), hasCredential (Q.Med, Q.Arb — both, since 2026-08-29), sameAs (LinkedIn), email, image. Emitted on /about/ only: memberOf — the four §4 memberships as Organization nodes (Q53, ruled 2026-08-28). / shows no memberships, so its Person node omits it: structured data represents the page it sits on. Withheld: worksFor — Q49(b) declined the row 2026-08-28 and Pouya confirmed the reading 2026-08-29, so it is settled rather than pending; provider → Person → worksFor would assert a same-entity claim §4 does not row. (This enumeration listed worksFor as emitted while the same cell said it was withheld, and omitted url and email, which are — wrong in both directions. The enumeration is the part an implementer copies. Found by adversarial-reviewer.) CHANGED 2026-08-28 — Q47. This row read "jobTitle = 'Director of Firm Operations'; omit worksFor", which put the boutique title on a node whose url is this ADR practice's /about/ — so a consumer could attach it to this entity. Pouya's ruling reframes the field: jobTitle describes this practice, not the boutique role, which D16 keeps unnamed. The visible role line is unchanged and still reads "Director of Firm Operations at a Toronto litigation and ADR boutique". THE VALUE IS PRACTICE_JOB_TITLE IN src/data/site.ts AND THIS ROW DOES NOT RESTATE IT — §7's rule, applied to a string with a live revert trigger on it: this row carried the literal text for one pass, and adversarial-reviewer noted it would go stale the moment the constant moved. Cite, do not copy. worksFor IS WITHHELD — set for one pass under Q47, then reverted: ProfessionalService.provider is this Person, so provider → Person → worksFor asserts the same-entity claim schema.ts explicitly declines, and §4 says "alongside the practice" where the ruling says "operates through". memberOf is emitted — see the sentence above; Q53 closed 2026-08-28. (This cell asserted memberOf was both emitted and withheld for one pass, which is the defect it already records itself being caught for on worksFor, in the opposite direction. The enumeration is the part an implementer copies.) See src/data/schema.ts
ProfessionalService Home areaServed Toronto/Ontario, serviceType Mediation / Commercial arbitration / Mediation-arbitration (med-arb)scoped 2026-08-28 on claims-auditor's finding; this row instructed the unscoped class form "Mediation/Arbitration" that Q39 struck and that schema.ts deliberately does not follow. Family arbitration carries prescribed training and has its own NOT OFFERED row, so unscoped "Arbitration" is the struck universal in a field nobody reads. Do not widen these strings without a §4 row to widen them fromprovider → Person, priceRange once /fees/ is real. Never LegalService — schema.org defines it as a business providing legal advice and representation, which asserts in machine-readable form exactly what D13 bars and §4 Forbidden calls out
Service /mediation/, /arbitration/, /med-arb/ and each practice page serviceType, provider → Person, areaServed. The Person node travels in the same @graph so provider: {'@id'} resolves in one document rather than relying on a crawler joining two — homeGraph's reasoning, applied. serviceType is scoped where §4 scopes it: Commercial arbitration, never a bare "Arbitration". No BreadcrumbList on the three — one hop from the root, no visible breadcrumb, and this spec requires the markup to match the visible one
Article Each article headline, description, datePublished, dateModified, author → Person, image
BreadcrumbList All nested pages Matches visible breadcrumbs
FAQPage /for-parties/, /med-arb/ Only where the visible page genuinely is Q&A. Never fabricate questions to farm a rich result

hasCredential must reflect reality. ⚠️ Q.Med AND Q.Arb are both held as of 2026-08-29 (AGENTS.md §4) and the field carries both — it was Q.Med-only while Q.Arb was a commenced pathway. The rule is unchanged and cuts both ways: marking an unheld credential as held is a misrepresentation that happens to be machine-readable, and omitting a held one understates the record in a field whose whole meaning is "holds". personNode maps CREDENTIALS.designations rather than indexing it, so a designation added to that constant reaches the graph automatically. That links the CONSTANT to the graph, not §4 to the graph — a designation added to §4 and not to src/data/site.ts still drifts silently, and no mechanism catches it. src/data/schema.ts states the condition in full; this sentence overstated it for one pass.

Crawlability

public/robots.txt is the artefact — read it, do not read a copy of it here. This spec used to reproduce the file inline and the reproduction had already drifted from it by 2026-08-26, which is the failure mode the AGENTS.md §7 rule exists to stop.

It disallows nothing, and that is deliberate. This spec previously prescribed Disallow: /legal/ alongside noindex on those pages, and the two cancel each other: a crawler forbidden to fetch a URL never reads the noindex on it. /legal/privacy/ and /legal/terms/ are linked from the footer of every page, so they are discovered regardless — and the likely result of the pair was Google listing the bare URLs as "no information available", the opposite of the intent, with the directive that would have suppressed them sitting unread behind the wall. noindex is what de-indexes; Disallow is what prevents fetching. Use the one that matches the problem, and never both on the same path.

Do not block AI crawlers. Being read by an assistant that a general counsel is using to shortlist neutrals is the point.

Sitemap: @astrojs/sitemap, excluding /legal/* and any draft: true article. Submit to Google Search Console and Bing Webmaster Tools at cutover.

Internal linking. Every practice page links to /mediation/ and /arbitration/; those link back to the practice areas; every article links to at least one practice page. This is what turns Insights into ranking power for the pages that convert. Breadcrumbs on every nested page.

404 page. Real, styled, with search-intent links out. CloudFront must return it with a genuine 404 status — not a 200, which the S3 website-endpoint pattern gets wrong by default.

Performance

Core Web Vitals are a ranking input, and the current build fails all of them.

Metric Budget
LCP < 2.0 s, Slow 4G
CLS < 0.05
INP < 150 ms
JS per route < 100 KB
Lighthouse (mobile) ≥ 95 all four categories — measurable again as of 2026-08-31, see below

THE INSTRUMENT IS BACK — build step 7a, 2026-08-31. npm run lighthouse, and it is lighthouse rather than @lhci/cli. R11's re-add trigger said to put @lhci/cli back; this is a deliberate deviation from its literal wording and AGENTS.md §7 records both the reason and what it costs.

The reason is that §7's advisory attribution was wrong, and it was the attribution that made the tool look unusable. §7 recorded the ten findings as arriving "via lighthouse → puppeteer-core → extract-zip". Measured from two probe lockfiles: @lhci/cli@0.15.1 carries 10 (7 high) and pins lighthouse 12.6.1, and the two high carriers are tmp@0.1.0its own direct dependency — and extract-zip@2.0.1 via @puppeteer/browsers. lighthouse@13.4.1 standalone is 109 packages, and both are absent: npm audit returns 0. So Lighthouse was never the carrier, and the budget was unmeasurable for five days on a cause nobody re-derived.

What it does not do: run in CI. Standalone Lighthouse drives an installed browser and the Gitea runner has none (§7, Q23). So it is a local gate plus a blocking item on docs/06's cutover checklist, and it is deliberately not wired into npm run build or either deploy path — a check described as running where it cannot is the defect Q22 turned out to be.

⚠️ THE ACCESSIBILITY CATEGORY IS MEASURED WITH prefers-reduced-motion FORCED, and that is a deviation that has to travel with the number. Measured twice per condition on /process/: motion on gives 96 with color-contrast failing on 24 nodes; motion off gives 100 with 0. The 24 were the scroll-driven reveal caught mid-flight — axe reported foregrounds like #d0cbc4 on #f8f4ed, and neither is in this palette; they are the real colours blended toward the background by an in-progress opacity keyframe. A category reporting 24 known-false nodes on ten of fourteen pages cannot surface the twenty-fifth real one. The reduced-motion rendering is the branch global.css ships for a real user setting, and it is the one where every element sits at its final colour. Palette ratios are computed in docs/02-design-system.md; scripts/lighthouse.mjs carries the measurement.

How: static HTML, self-hosted preloaded subset fonts, AVIF/WebP with explicit dimensions, critical CSS inlined, no third-party scripts on any page except the booking embed on /contact/ — and that one is lazy-loaded behind a click.

Local and professional presence

Not code, but it belongs in the launch checklist: Google Business Profile for the practice; ADRIC and ADRIO directory listings pointing at the site; a LinkedIn profile whose headline and Featured section match the brand (brief §VIII); consistent naming across all of them. Not "name, address and phone" — §4 publishes no phone number and no street address, only "Toronto, Ontario; by appointment". Directory forms that demand a full NAP get what §4 verifies and nothing more.

Post-launch verification

  • curl -s https://adr.smlcompany.ca/ | grep -c "<h1" returns ≥ 1
  • Every page renders its full text with JavaScript disabled
  • Rich Results Test passes on Person, ProfessionalService, Article
  • OG preview renders correctly in LinkedIn Post Inspector and Slack
  • Sitemap submitted to Google Search Console and Bing
  • No page returns 200 for a URL that should 404
  • Lighthouse ≥ 95 mobile on every built pagenpm run lighthouse, which enumerates dist/ rather than taking a list, so the set cannot go stale as pages are added. Do not tick this box from a manual Chrome DevTools run and call it the same check