Files
adr-sml/docs/04-seo-spec.md
T
Pouya LajevardiandClaude Opus 5 8134709548 feat: build step 1 — scaffold, layout, header, footer, SEO; zero JavaScript
Build order step 1 (docs/01): scaffold, tokens, base layout, header,
footer, SEO component, plus a temporary /type-scale/ proof sheet that
step 2 deletes.

THE FONTS WERE NEVER ON DISK. global.css declared six @font-face rules
pointing at /fonts/*.woff2 and public/fonts/ did not exist, so every
face had been silently falling back to Georgia and the system sans.
Six cuts committed, 123,804 bytes, SIL OFL 1.1, provenance in
docs/reference/fonts-provenance.md. ?v=1 on every URL because the
deploy script serves them immutable for a year.

ZERO JAVASCRIPT. The reveal was an inline IntersectionObserver in
<head>; docs/05 specifies script-src 'self' with no unsafe-inline, so
the only script on the site was the one thing the site's own CSP would
refuse to execute. Replaced with animation-timeline: view() behind
@supports. 0 script tags and 0 .js files in dist.

The infinity mark is lifted verbatim from the deployed site's own
smlMark loading thumbnail, not redrawn (Q32 asks whether a canonical
vector exists). The proof sheet computes its contrast table from
tokens.css rather than restating docs/02 — all eleven ratios reproduce
the measured table exactly.

Register: Canadian Tax Foundation added (§4, R10 widened); Q30 closed
— SML Company Ltd is federally incorporated under the CBCA, and the
footer publishes neither that nor the place of business; Q31 closed —
Plausible, on EU-only data residency (D15 amended). ROLE constants
added for "Director of Firm Operations" and "active litigation
exposure" so step 3 does not hand-type them.

Lighthouse unavailability now stated in six places rather than left as
a control that had silently stopped existing (§7, R11).

Both review agents ran twice. The second pass found four defects in
the first pass's fixes, including the minifier bug written back into
its own fix and a colour-alone repair that used the banned gold-on-
cream pairing at 2.10:1. Measured in headless Chrome at thirteen
widths with a seventh nav item injected: 0 overflow, 0 tap targets
under 44x44, 0 focus-order inversions, state indicators at 12.29:1,
755 words of body text with no JavaScript.

Opened: Q32-Q37. Closed: Q30, Q31.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012XquaEq4BgWMCwUqLEyNkF
2026-08-26 15:57:02 -04:00

169 lines
8.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 04 — Discoverability
The problem this project exists to fix. `AGENTS.md` §2 has the measurements.
> **Those measurements are under review — `AGENTS.md` Q34.** They were taken on
> 2026-08-25; a re-fetch on 2026-08-26 returned a bundler harness whose real
> `<head>` sits JSON-escaped inside a `<script>` and whose application lives in
> nine UUID-named files that were not fetched. Some of §2 reproduced exactly
> (the 2.2 MB single file, the placeholder `<title>`); some could not be
> reproduced from the served HTML at all. **Cite §2, and cite Q34 with it. Do
> not put any of these figures in public copy until Q34 closes.**
---
## The baseline being replaced
| | Now `[verified 2026-08-25]` | Target |
|---|---|---|
| Content in server HTML | `SML Company · DISPUTE RESOLUTION · Unpacking...` | Every word |
| Indexable pages | 1 | 17 + articles (19 fixed URLs, less the two `/legal/*` pages, which are `noindex` and excluded from the sitemap) |
| `<title>` | `SML Company · Dispute Resolution` — pre-rebrand placeholder | Unique per page |
| Meta description | none | Unique per page |
| `<meta viewport>` | **absent** | Present |
| Canonical URL | none | Every page |
| OG / Twitter tags | none | Every page |
| Structured data | none | Person, ProfessionalService, Article, FAQ, Breadcrumb |
| `robots.txt` | 403 | Served |
| Sitemap | none | Generated at build |
| Favicon | none | Full set |
Astro's static output solves most of this by existing. The rest is the spec below.
## Why static output matters more than usual here
Google can sometimes render client-side JavaScript. Bing largely does not.
LinkedIn's preview crawler does not. Slack's unfurler does not. **And the
crawlers behind AI assistants — increasingly how counsel and in-house teams
find a neutral — generally do not.**
A site that requires three CDN round trips and an in-browser Babel compile before
producing a sentence is invisible to all of them. That is the whole argument for
D1.
---
## Metadata
Every page passes through one `SEO` component. A page without it is not finished.
```
title 5060 chars, unique. Pattern: "<Page> · Pouya Lajevardi"
Home: "Pouya Lajevardi · Mediation & Arbitration · Toronto"
ARTICLES ARE THE EXCEPTION: no " · Pouya Lajevardi" suffix.
The suffix is 18 chars, so a headline that already reads
5060 renders at 6878 — over this ceiling. Measured against
the five launch headlines in 03-content-spec.md, the suffix
rule fails 5 of 5; without it, 4 of 5 pass. An article's
headline IS its <title>; `seoTitle` in the frontmatter
overrides it when a headline that reads well is out of range.
src/content.config.ts enforces this and names the offending
string and its length in the build error.
description 140160 chars, unique, written for a human, not stuffed
canonical absolute, https, trailing slash
og:title/description/image/url/type/site_name/locale (en_CA)
twitter:card summary_large_image
robots index,follow — except /legal/* which is noindex,follow
```
**OG images:** 1200 × 630. Generate at build with `satori` or `astro-og-canvas`
using the site's own type and palette. One template: display headline on cream,
infinity mark, designation line. Never a screenshot.
## Structured data
JSON-LD only. Validate against Google's Rich Results Test before cutover.
| Type | Where | Notes |
|---|---|---|
| `Person` | `/about/`, referenced site-wide | `name`, `jobTitle`, `description`, `alumniOf` (Bond University), `knowsLanguage` (en, fa), `hasCredential` (Q.Med), `sameAs` (LinkedIn), `image`. **`jobTitle` = "Director of Firm Operations"; omit `worksFor`** — populating it either names the boutique (D16) or misstates the employer |
| `ProfessionalService` | Home | `areaServed` Toronto/Ontario, `serviceType` Mediation/Arbitration, `provider` → Person, `priceRange` once `/fees/` is real. **Never `LegalService`** — schema.org defines it as a business providing legal advice and *representation*, which asserts in machine-readable form exactly what D13 bars and §4 Forbidden calls out |
| `Service` | Each practice page | `serviceType`, `provider` → Person, `areaServed` |
| `Article` | Each article | `headline`, `description`, `datePublished`, `dateModified`, `author` → Person, `image` |
| `BreadcrumbList` | All nested pages | Matches visible breadcrumbs |
| `FAQPage` | `/for-parties/`, `/med-arb/` | Only where the visible page genuinely is Q&A. Never fabricate questions to farm a rich result |
**`hasCredential` must reflect reality.** Q.Med is held. Q.Arb is not. Marking an
unheld credential as held in structured data is a misrepresentation that happens
to be machine-readable.
## Crawlability
**`public/robots.txt` is the artefact — read it, do not read a copy of it
here.** This spec used to reproduce the file inline and the reproduction had
already drifted from it by 2026-08-26, which is the failure mode the `AGENTS.md`
§7 rule exists to stop.
**It disallows nothing, and that is deliberate.** This spec previously
prescribed `Disallow: /legal/` alongside `noindex` on those pages, and the two
cancel each other: a crawler forbidden to *fetch* a URL never reads the
`noindex` on it. `/legal/privacy/` and `/legal/terms/` are linked from the
footer of every page, so they are discovered regardless — and the likely result
of the pair was Google listing the bare URLs as "no information available", the
opposite of the intent, with the directive that would have suppressed them
sitting unread behind the wall. **`noindex` is what de-indexes; `Disallow` is
what prevents fetching.** Use the one that matches the problem, and never both
on the same path.
Do not block AI crawlers. Being read by an assistant that a general counsel is
using to shortlist neutrals is the point.
**Sitemap:** `@astrojs/sitemap`, excluding `/legal/*` and any `draft: true`
article. Submit to Google Search Console and Bing Webmaster Tools at cutover.
**Internal linking.** Every practice page links to `/mediation/` and
`/arbitration/`; those link back to the practice areas; every article links to
at least one practice page. This is what turns Insights into ranking power for
the pages that convert. Breadcrumbs on every nested page.
**404 page.** Real, styled, with search-intent links out. CloudFront must return
it with a genuine 404 status — not a 200, which the S3 website-endpoint pattern
gets wrong by default.
## Performance
Core Web Vitals are a ranking input, and the current build fails all of them.
| Metric | Budget |
|---|---|
| LCP | < 2.0 s, Slow 4G |
| CLS | < 0.05 |
| INP | < 150 ms |
| JS per route | < 100 KB |
| Lighthouse (mobile) | ≥ 95 all four categories — **not measurable until step 7, see below** |
> ⚠️ **Lighthouse verification is UNAVAILABLE until build step 7.** `@lhci/cli`
> was removed on 2026-08-26 — it was the sole source of all 10 `npm audit`
> findings (7 high), `0.15.1` is `latest` so there was no clean upgrade, and it
> could not run at all with no pages and no `lighthouserc`. The budget below is
> not suspended; the tool that measures it is absent. Re-add at step 7 under
> `AGENTS.md` R11, checking for a patched release rather than assuming `0.15.1`
> is still the ceiling. Until then, a run that skips this is skipping something
> known — not something forgotten. `AGENTS.md` §7 has the state.
How: static HTML, self-hosted preloaded subset fonts, AVIF/WebP with explicit
dimensions, critical CSS inlined, no third-party scripts on any page except the
booking embed on `/contact/` — and that one is lazy-loaded behind a click.
## Local and professional presence
Not code, but it belongs in the launch checklist: Google Business Profile for the
practice; ADRIC and ADRIO directory listings pointing at the site; a LinkedIn
profile whose headline and Featured section match the brand (brief §VIII);
consistent naming across all of them. **Not "name, address and phone"** — §4
publishes no phone number and no street address, only "Toronto, Ontario; by
appointment". Directory forms that demand a full NAP get what §4 verifies and
nothing more.
## Post-launch verification
- [ ] `curl -s https://adr.smlcompany.ca/ | grep -c "<h1"` returns ≥ 1
- [ ] Every page renders its full text with JavaScript disabled
- [ ] Rich Results Test passes on Person, ProfessionalService, Article
- [ ] OG preview renders correctly in LinkedIn Post Inspector and Slack
- [ ] Sitemap submitted to Google Search Console and Bing
- [ ] No page returns 200 for a URL that should 404
- [ ] Lighthouse ≥ 95 mobile on `/`, `/about/`, one practice page, one article
**blocked until `@lhci/cli` is re-added at step 7.** Do not tick this box
from a manual Chrome DevTools run and call it the same check