A security-first FastAPI service for technical SEO, SaaS acquisition strategy, sampled site crawling, page comparison, and first-party opportunity ranking.
This fork is a ground-up v2 implementation of KovalDenys1/SEO-Analyzer-API. It keeps the original GET endpoints for compatibility while replacing the one-file analyzer with a tested, explainable, asynchronous engine.
- Audits status, redirects, indexability, canonical hints, metadata, headings, content, internal links, images, language, mobile setup, social previews, and JSON-LD.
- Classifies SaaS pages such as pricing, feature, use case, industry, integration, template, comparison, alternative, free tool, docs, case study, security, and changelog pages.
- Scores core SEO health separately from SaaS acquisition/conversion readiness.
- Produces evidence-backed issues and ranked recommendations with impact, effort, confidence, validation, and affected pages.
- Crawls a bounded site sample while respecting
robots.txt, discovering XML sitemap indexes, and detecting sampled broken links, exact duplicates, weak internal linking, and potential orphan pages. - Runs persistent unified link-graph scans: one bounded crawl builds internal/external link graph data and attaches SEO results to every graph node.
- Assesses seven SaaS strategy pillars: commercial foundation, audience/use cases, product-led acquisition, bottom-funnel evaluation, authority, trust, and product enablement.
- Compares up to eight pages without mislabeling the result as a Google ranking.
- Ranks supplied Search Console/conversion rows by traffic and revenue opportunity without inventing external keyword or SERP data.
- Optionally enriches a page with Google PageSpeed Insights v5 lab/field data.
flowchart LR
Client[API or browser UI] --> API[FastAPI]
API --> Jobs[Bounded scan manager]
Jobs --> Crawler[Shared SiteCrawler]
Crawler --> Fetcher[SafeFetcher]
Crawler --> Parser[HTML parser and SEO scoring]
Parser --> Graph[Link graph and site-level issues]
Graph --> Storage[(SQLite)]
Storage --> API
API --> Dashboard[Interactive graph dashboard]
/v1/site-audit and persistent link-graph scans use the same crawler engine.
Each fetched document is parsed once, then reused for SEO scoring, internal and
external edges, duplicate/orphan/broken-link checks, storage, and visualization.
The service fetches user-supplied URLs, so URL handling is part of the security boundary:
- only HTTP(S) on configured ports;
- credentials in URLs are rejected;
- private, loopback, link-local, multicast, and reserved addresses are rejected;
- every DNS answer and every redirect target is validated;
- the request connects to the validated IP while preserving the public Host/SNI identity, limiting DNS-rebinding exposure;
- proxy environment variables are ignored;
- response time, redirect count, concurrency, and decompressed response size are bounded;
- optional
X-API-Keyauthentication; - bounded TTL/LRU cache and request coalescing.
Private-network access can be enabled for a trusted internal deployment, but it is off by default.
| Method | Path | Purpose |
|---|---|---|
GET |
/v1/analyze?url=… |
Full page report; optional PageSpeed and subdomain scope |
POST |
/v1/site-audit |
Bounded robots-aware crawl and SaaS strategy assessment |
POST |
/v1/compare |
Relative SEO/SaaS comparison for 2–8 pages |
POST |
/v1/opportunities |
First-party traffic/revenue opportunity ranking |
GET |
/app |
Minimal browser UI for unified link-graph scans |
POST |
/api/projects |
Create a persistent crawl project |
GET |
/api/projects |
List projects |
GET |
/api/projects/{project_id} |
Get one project |
POST |
/api/projects/{project_id}/scans |
Start a background crawl + graph + SEO scan |
GET |
/api/projects/{project_id}/scans |
List project scans |
GET |
/api/scans/{scan_id}/status |
Get progress and terminal status |
POST |
/api/scans/{scan_id}/cancel |
Cancel pending or running work |
POST |
/api/scans/{scan_id}/rerun |
Create a new scan with the same options |
GET |
/api/scans/{scan_id}/pages |
List page records and attached SEO data |
GET |
/api/scans/{scan_id}/page?url=... |
Resolve a page by normalized URL |
GET |
/api/scans/{scan_id}/pages/{graph_node_id} |
Resolve a page by graph node ID |
GET |
/api/scans/{scan_id}/links |
List internal, external, or redirect links |
GET |
/api/scans/{scan_id}/graph |
Graph nodes/edges with SEO data attached |
GET |
/api/scans/{scan_id}/seo/issues |
List page and site-level SEO issues |
GET |
/api/scans/{scan_id}/stats |
Site totals, duplicates, cycles, orphans, and failures |
GET |
/api/scans/{scan_id}/dashboard |
Interactive graph dashboard for a completed scan |
GET |
/analyze |
Backwards-compatible original full-analysis shape plus v2 data |
GET |
/quick-score |
Score, warnings, and top recommendations |
GET |
/metadata |
Metadata, headings, social, and canonical data |
GET |
/healthz, /readyz, /metrics |
Health, readiness, and Prometheus metrics |
Interactive OpenAPI docs are available at /docs and /redoc.
Python 3.11–3.14 is supported.
python -m venv .venv
. .venv/bin/activate
python -m pip install -e '.[dev]'
uvicorn main:app --reloadAnalyze a page:
curl --get http://127.0.0.1:8000/v1/analyze \
--data-urlencode 'url=https://example.com'Audit a site sample:
curl http://127.0.0.1:8000/v1/site-audit \
-H 'Content-Type: application/json' \
-d '{
"url": "https://example.com",
"max_pages": 25,
"max_depth": 3,
"concurrency": 5,
"respect_robots": true,
"use_sitemap": true
}'Run a persistent link-graph scan:
curl -X POST http://127.0.0.1:8000/api/projects \
-H 'Content-Type: application/json' \
-d '{"url":"https://example.com","name":"Example"}'
curl -X POST http://127.0.0.1:8000/api/projects/{project_id}/scans \
-H 'Content-Type: application/json' \
-d '{"max_pages":50,"max_depth":4,"concurrency":4,"respect_robots":true}'
curl http://127.0.0.1:8000/api/scans/{scan_id}/status
curl http://127.0.0.1:8000/api/scans/{scan_id}/graphThe browser workflow is available at http://127.0.0.1:8000/app.
If SEO_API_KEY is set, add -H 'X-API-Key: …' to protected endpoints.
The image runs as a non-root user. The Compose example binds only to loopback,
drops Linux capabilities, uses a read-only root filesystem, and persists SQLite
in the analyzer-data volume mounted at /data.
cp .env.example .env
docker compose up --buildAll settings use the SEO_ prefix. See .env.example for the complete list.
| Variable | Default | Meaning |
|---|---|---|
SEO_API_KEY |
empty | Optional X-API-Key shared secret |
SEO_FETCH_TIMEOUT_SECONDS |
12 |
Upstream request timeout |
SEO_MAX_RESPONSE_BYTES |
3000000 |
Maximum decompressed page/resource body |
SEO_MAX_REDIRECTS |
5 |
Redirect budget |
SEO_MAX_CONCURRENT_FETCHES |
8 |
Process-wide fetch concurrency |
SEO_ALLOW_PRIVATE_HOSTS |
false |
Permit non-public targets; trusted deployments only |
SEO_CACHE_TTL_SECONDS |
300 |
Analysis cache TTL; 0 disables cache |
SEO_MAX_SITE_PAGES |
100 |
Server-side hard cap for a site audit |
SEO_SCAN_STORAGE_PATH |
data/analyzer.db |
SQLite storage for persistent link-graph scans |
SEO_SCAN_JOB_WORKERS |
2 |
Concurrent link-graph scans per API process |
SEO_SCAN_JOB_LEASE_SECONDS |
30 |
Time before an abandoned running scan is requeued |
SEO_ENABLE_PAGESPEED |
false |
Permit quota-consuming PageSpeed calls |
SEO_PAGESPEED_API_KEY |
empty | Optional Google API key |
SEO_CORS_ORIGINS |
empty | Comma-separated browser origins |
The core score is a weighted, fully explainable health summary. Every deduction maps to an issue code and evidence. The SaaS score is a separate page-type-aware acquisition/conversion heuristic. Site strategy maturity measures detected coverage in the bounded sample.
None of these scores predicts a Google position, traffic, revenue, content quality, or rich-result eligibility. The network timer is not LCP/INP/CLS. Browser performance is reported only when PageSpeed is explicitly enabled. See the scoring methodology for weights and caveats.
ruff check .
ruff format --check .
mypy seo_analyzer
pytest --cov --cov-report=term-missing
pip-auditThe suite covers URL security, DNS/redirect validation, parsing, page classification, issue scoring, sitemap/robots behavior, crawl aggregation, PageSpeed normalization, API compatibility, auth, and opportunity ranking.
SQLite stores projects, scans, normalized page URLs, stable graph node IDs, links with page zones, SEO reports, status/depth/timing metadata, and the complete versioned result snapshot. HTML response bodies are not persisted. Running scans write a worker lease and heartbeat; graceful shutdown requeues active work, while an abandoned lease is recovered automatically.
- Crawls are bounded samples and do not execute client-side JavaScript.
- External targets are represented in the graph but are not fetched or checked.
- SQLite is intended for local and small-team deployments; high-volume distributed deployments should replace the storage/job backend.
- Worker limits apply per API process. Global fetch concurrency is still bounded
by
SEO_MAX_CONCURRENT_FETCHESin each process.
- API guide
- Scoring methodology
- SaaS SEO strategy playbook
- Original-project audit
- Security model
- Changelog
MIT. The upstream project is copyright Denys Koval. Link-graph functionality is derived from ruslan2027/link-graph-explorer-oss. See LICENSE, THIRD_PARTY_NOTICES.md, and licenses/link-graph-explorer-oss.LICENSE.