The re-audit of the deploy-guard change surfaced defects well outside the diff, including one that would have broken production mail. docs/05-backend-spec.md had the two SES DKIM sets exactly inverted, labelling the three records that resolve as "orphans" and the three NXDOMAIN records as "Live. Never delete". Entry (j) corrected this in AGENTS.md §7 and the correction never reached docs/05. Since SES has no custom MAIL FROM, DKIM is the only thing satisfying DMARC, so acting on that table would have silently broken intake mail authentication. Also in this change: - .gitea/workflows/deploy.yml gains a guard as steps[0] that fails the run, naming the variable, if AWS_REGION, S3_BUCKET or CLOUDFRONT_DISTRIBUTION_ID is empty — how a Gitea too old for the vars context manifests. Verified fail-closed under bash -e, sh -e and bash -euo pipefail. - AGENTS.md Current Truth: SPF and DMARC recorded as present (Q20), the matching §10 High risk row retired, three duplicate Q rows removed. - docs/reference/AWS-Hosting-Guide.md tracked and given a do-not-execute banner; it was an executable procedure for the architecture D1/D3 replace. - Copy decks: "a working litigator" and "an active litigation practice" replaced with the register's own wording; LegalService JSON-LD replaced with ProfessionalService; tribunal-secretary offers removed per D14; nine stale question blockers swept. - astro.config.mjs: prefetchAll disabled — it injected JS into every page against the zero-JS convention with no decision recorded. - src/data/site.ts: unregistered response-time commitment nulled (Q27); OBA section names downgraded to [assumed] (Q28). - s3:AbortMultipartUpload reasoning corrected to measure ./dist, not the repo. Opens Q27, Q28, Q29. AGENTS.md entry (q) records the full resolution, including the findings declined and why. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012XquaEq4BgWMCwUqLEyNkF
5.8 KiB
04 — Discoverability
The problem this project exists to fix. AGENTS.md §2 has the measurements: a
server-side fetch of the live site returns three words.
The baseline being replaced
Now [verified 2026-08-25] |
Target | |
|---|---|---|
| Content in server HTML | SML Company · DISPUTE RESOLUTION · Unpacking... |
Every word |
| Indexable pages | 1 | 17 + articles (19 fixed URLs, less the two /legal/* pages, which are noindex and excluded from the sitemap) |
<title> |
SML Company · Dispute Resolution — pre-rebrand placeholder |
Unique per page |
| Meta description | none | Unique per page |
<meta viewport> |
absent | Present |
| Canonical URL | none | Every page |
| OG / Twitter tags | none | Every page |
| Structured data | none | Person, ProfessionalService, Article, FAQ, Breadcrumb |
robots.txt |
403 | Served |
| Sitemap | none | Generated at build |
| Favicon | none | Full set |
Astro's static output solves most of this by existing. The rest is the spec below.
Why static output matters more than usual here
Google can sometimes render client-side JavaScript. Bing largely does not. LinkedIn's preview crawler does not. Slack's unfurler does not. And the crawlers behind AI assistants — increasingly how counsel and in-house teams find a neutral — generally do not.
A site that requires three CDN round trips and an in-browser Babel compile before producing a sentence is invisible to all of them. That is the whole argument for D1.
Metadata
Every page passes through one SEO component. A page without it is not finished.
title 50–60 chars, unique. Pattern: "<Page> · Pouya Lajevardi"
Home: "Pouya Lajevardi · Mediation & Arbitration · Toronto"
description 140–160 chars, unique, written for a human, not stuffed
canonical absolute, https, trailing slash
og:title/description/image/url/type/site_name/locale (en_CA)
twitter:card summary_large_image
robots index,follow — except /legal/* which is noindex,follow
OG images: 1200 × 630. Generate at build with satori or astro-og-canvas
using the site's own type and palette. One template: display headline on cream,
infinity mark, designation line. Never a screenshot.
Structured data
JSON-LD only. Validate against Google's Rich Results Test before cutover.
| Type | Where | Notes |
|---|---|---|
Person |
/about/, referenced site-wide |
name, jobTitle, description, alumniOf (Bond University), knowsLanguage (en, fa), hasCredential (Q.Med), sameAs (LinkedIn), image. jobTitle = "Director of Firm Operations"; omit worksFor — populating it either names the boutique (D16) or misstates the employer |
ProfessionalService |
Home | areaServed Toronto/Ontario, serviceType Mediation/Arbitration, provider → Person, priceRange once /fees/ is real. Never LegalService — schema.org defines it as a business providing legal advice and representation, which asserts in machine-readable form exactly what D13 bars and §4 Forbidden calls out |
Service |
Each practice page | serviceType, provider → Person, areaServed |
Article |
Each article | headline, description, datePublished, dateModified, author → Person, image |
BreadcrumbList |
All nested pages | Matches visible breadcrumbs |
FAQPage |
/for-parties/, /med-arb/ |
Only where the visible page genuinely is Q&A. Never fabricate questions to farm a rich result |
hasCredential must reflect reality. Q.Med is held. Q.Arb is not. Marking an
unheld credential as held in structured data is a misrepresentation that happens
to be machine-readable.
Crawlability
public/robots.txt:
User-agent: *
Allow: /
Disallow: /legal/
Sitemap: https://adr.smlcompany.ca/sitemap-index.xml
Do not block AI crawlers. Being read by an assistant that a general counsel is using to shortlist neutrals is the point.
Sitemap: @astrojs/sitemap, excluding /legal/* and any draft: true
article. Submit to Google Search Console and Bing Webmaster Tools at cutover.
Internal linking. Every practice page links to /mediation/ and
/arbitration/; those link back to the practice areas; every article links to
at least one practice page. This is what turns Insights into ranking power for
the pages that convert. Breadcrumbs on every nested page.
404 page. Real, styled, with search-intent links out. CloudFront must return it with a genuine 404 status — not a 200, which the S3 website-endpoint pattern gets wrong by default.
Performance
Core Web Vitals are a ranking input, and the current build fails all of them.
| Metric | Budget |
|---|---|
| LCP | < 2.0 s, Slow 4G |
| CLS | < 0.05 |
| INP | < 150 ms |
| JS per route | < 100 KB |
| Lighthouse (mobile) | ≥ 95 all four categories |
How: static HTML, self-hosted preloaded subset fonts, AVIF/WebP with explicit
dimensions, critical CSS inlined, no third-party scripts on any page except the
booking embed on /contact/ — and that one is lazy-loaded behind a click.
Local and professional presence
Not code, but it belongs in the launch checklist: Google Business Profile for the practice; ADRIC and ADRIO directory listings pointing at the site; a LinkedIn profile whose headline and Featured section match the brand (brief §VIII); consistent name, address, and phone across all of them.
Post-launch verification
curl -s https://adr.smlcompany.ca/ | grep -c "<h1"returns ≥ 1- Every page renders its full text with JavaScript disabled
- Rich Results Test passes on Person, ProfessionalService, Article
- OG preview renders correctly in LinkedIn Post Inspector and Slack
- Sitemap submitted to Google Search Console and Bing
- No page returns 200 for a URL that should 404
- Lighthouse ≥ 95 mobile on
/,/about/, one practice page, one article