Files
adr-sml/docs/04-seo-spec.md
T
Pouya Lajevardi 19f7226661
Build and deploy / build-and-deploy (push) Failing after 5s
chore: project scaffold, specs, and working record
2026-08-26 08:51:16 -04:00

133 lines
5.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 04 — Discoverability
The problem this project exists to fix. `AGENTS.md` §2 has the measurements: a
server-side fetch of the live site returns three words.
---
## The baseline being replaced
| | Now `[verified 2026-08-25]` | Target |
|---|---|---|
| Content in server HTML | `SML Company · DISPUTE RESOLUTION · Unpacking...` | Every word |
| Indexable pages | 1 | 19 + articles |
| `<title>` | `SML Company · Dispute Resolution` — pre-rebrand placeholder | Unique per page |
| Meta description | none | Unique per page |
| `<meta viewport>` | **absent** | Present |
| Canonical URL | none | Every page |
| OG / Twitter tags | none | Every page |
| Structured data | none | Person, LegalService, Article, FAQ, Breadcrumb |
| `robots.txt` | 403 | Served |
| Sitemap | none | Generated at build |
| Favicon | none | Full set |
Astro's static output solves most of this by existing. The rest is the spec below.
## Why static output matters more than usual here
Google can sometimes render client-side JavaScript. Bing largely does not.
LinkedIn's preview crawler does not. Slack's unfurler does not. **And the
crawlers behind AI assistants — increasingly how counsel and in-house teams
find a neutral — generally do not.**
A site that requires three CDN round trips and an in-browser Babel compile before
producing a sentence is invisible to all of them. That is the whole argument for
D1.
---
## Metadata
Every page passes through one `SEO` component. A page without it is not finished.
```
title 5060 chars, unique. Pattern: "<Page> · Pouya Lajevardi"
Home: "Pouya Lajevardi · Mediation & Arbitration · Toronto"
description 140160 chars, unique, written for a human, not stuffed
canonical absolute, https, trailing slash
og:title/description/image/url/type/site_name/locale (en_CA)
twitter:card summary_large_image
robots index,follow — except /legal/* which is noindex,follow
```
**OG images:** 1200 × 630. Generate at build with `satori` or `astro-og-canvas`
using the site's own type and palette. One template: display headline on cream,
infinity mark, designation line. Never a screenshot.
## Structured data
JSON-LD only. Validate against Google's Rich Results Test before cutover.
| Type | Where | Notes |
|---|---|---|
| `Person` | `/about/`, referenced site-wide | `name`, `jobTitle`, `description`, `alumniOf` (Bond University), `knowsLanguage` (en, fa), `hasCredential` (Q.Med), `sameAs` (LinkedIn — **Q12**), `image`, `worksFor` |
| `LegalService` | Home | `areaServed` Toronto/Ontario, `serviceType` Mediation/Arbitration, `provider` → Person, `priceRange` once `/fees/` is real |
| `Service` | Each practice page | `serviceType`, `provider` → Person, `areaServed` |
| `Article` | Each article | `headline`, `description`, `datePublished`, `dateModified`, `author` → Person, `image` |
| `BreadcrumbList` | All nested pages | Matches visible breadcrumbs |
| `FAQPage` | `/for-parties/`, `/med-arb/` | Only where the visible page genuinely is Q&A. Never fabricate questions to farm a rich result |
**`hasCredential` must reflect reality.** Q.Med is held. Q.Arb is not. Marking an
unheld credential as held in structured data is a misrepresentation that happens
to be machine-readable.
## Crawlability
**`public/robots.txt`:**
```
User-agent: *
Allow: /
Disallow: /legal/
Sitemap: https://adr.smlcompany.ca/sitemap-index.xml
```
Do not block AI crawlers. Being read by an assistant that a general counsel is
using to shortlist neutrals is the point.
**Sitemap:** `@astrojs/sitemap`, excluding `/legal/*` and any `draft: true`
article. Submit to Google Search Console and Bing Webmaster Tools at cutover.
**Internal linking.** Every practice page links to `/mediation/` and
`/arbitration/`; those link back to the practice areas; every article links to
at least one practice page. This is what turns Insights into ranking power for
the pages that convert. Breadcrumbs on every nested page.
**404 page.** Real, styled, with search-intent links out. CloudFront must return
it with a genuine 404 status — not a 200, which the S3 website-endpoint pattern
gets wrong by default.
## Performance
Core Web Vitals are a ranking input, and the current build fails all of them.
| Metric | Budget |
|---|---|
| LCP | < 2.0 s, Slow 4G |
| CLS | < 0.05 |
| INP | < 150 ms |
| JS per route | < 100 KB |
| Lighthouse (mobile) | ≥ 95 all four categories |
How: static HTML, self-hosted preloaded subset fonts, AVIF/WebP with explicit
dimensions, critical CSS inlined, no third-party scripts on any page except the
booking embed on `/contact/` — and that one is lazy-loaded behind a click.
## Local and professional presence
Not code, but it belongs in the launch checklist: Google Business Profile for the
practice; ADRIC and ADRIO directory listings pointing at the site; a LinkedIn
profile whose headline and Featured section match the brand (brief §VIII);
consistent name, address, and phone across all of them.
## Post-launch verification
- [ ] `curl -s https://adr.smlcompany.ca/ | grep -c "<h1"` returns ≥ 1
- [ ] Every page renders its full text with JavaScript disabled
- [ ] Rich Results Test passes on Person, LegalService, Article
- [ ] OG preview renders correctly in LinkedIn Post Inspector and Slack
- [ ] Sitemap submitted to Google Search Console and Bing
- [ ] No page returns 200 for a URL that should 404
- [ ] Lighthouse ≥ 95 mobile on `/`, `/about/`, one practice page, one article