A web dashboard for monitoring and analyzing bot/crawler traffic across web projects. Tracks AI crawlers, search engines, SEO tools, social previews, and more — so you know who is reading your content and why.
- Bot Detection — Identifies 130+ bots by User-Agent across AI training crawlers, search engines, SEO tools, social platforms, monitoring services, and CLI tools
- Bot Verification — Confirms bot identity via reverse DNS (PTR) lookups and known IP CIDR ranges
- Multi-Project Support — Track bot traffic across multiple web projects from a single dashboard
- Trend Analysis — Period-over-period comparison (24h / 7d / 30d / 90d / 1y, or a custom date range) showing rising bots, pages, and projects
- Status Quality — Response-code rollups, top failing paths, API/sensitive path hits, and bots with error or UA-only traffic
- AI Crawler Intel — Dedicated view for AI training and search crawlers with confidence breakdowns (verified vs UA-only) and a crawls-vs-visits breakdown by company
- Raw Events — Filterable event log with bot name, path, project, IP, and user-agent details
- Data Health Monitoring — Heartbeat freshness tracking to detect logging pipeline issues
- Next.js 16 (App Router, Turbopack)
- TypeScript
- Tailwind CSS v4
- Recharts (chart components)
- PostgreSQL (Aiven or any standard Postgres)
- Node.js 20+
- A PostgreSQL database (Aiven, Neon, or any standard Postgres)
# Install dependencies
npm install
# Set environment variables
cp .env.example .env
# Edit .env with your values| Variable | Required | Description |
|---|---|---|
DATABASE_URL |
Yes | PostgreSQL connection string |
BOT_ADMIN_TOKEN |
Yes | Dashboard login and session-signing secret; use a unique 32+ character value |
BOT_IP_HASH_SECRET |
Yes | Dedicated 32+ character secret for keyed IP hashes; do not reuse another role's secret |
BOT_INGEST_TOKENS |
Yes | JSON object mapping project names to unique 32+ character ingestion keys |
Generate each secret with openssl rand -base64 32 or an equivalent cryptographically secure generator. During migration, the legacy BOT_LOG_TOKEN is accepted only when no new ingestion mapping is configured.
Migrations live in db/migrations/*.sql and are applied in order by scripts/migrate.mjs, which tracks what's already been applied in a schema_migrations table (safe to re-run):
npm run migrate
# or point it at a specific database:
node scripts/migrate.mjs "$DATABASE_URL"The bot_hits_daily and bot_first_seen tables are seeded from existing history by the migration and then kept current by insertHit on every event. Migration 004_weighted_rollups.sql repairs the daily rollup for existing deployments where sampled rows were previously backfilled with unweighted counts. If you apply migrations while an older (pre-rollup) build is still receiving traffic, those rows land in bot_hits but not the rollup; after deploying, run the reconcile script once to rebuild the rollup from raw and restore exact parity (safe to re-run any time you suspect drift):
npm run reconcile-rollups
# or: node scripts/reconcile-rollups.mjs "$DATABASE_URL"npm run devOpen http://localhost:3000. The landing page is at /, the dashboard at /dashboard.
The app is a standard Next.js app and works on Vercel or any Node host that can run next start.
For Vercel:
- Create a PostgreSQL database (Aiven, Neon, or any provider).
- Run
npm run migrate(ornode scripts/migrate.mjs "$DATABASE_URL"). - Generate separate admin, IP-hash, and per-project ingestion secrets.
- Add
DATABASE_URL,BOT_ADMIN_TOKEN,BOT_IP_HASH_SECRET, andBOT_INGEST_TOKENSas environment variables. - Deploy the repository.
- Open
/dashboardand sign in withBOT_ADMIN_TOKEN.
Do not expose any of these secrets in client-side browser code. Tracked sites should send events from a server-only proxy, route, function, or backend logger.
The dashboard has 4 tabs, selected via ?view=:
| Route | Description |
|---|---|
/ |
Landing page with feature overview and login |
/dashboard or /dashboard?view=overview |
Totals, crawler mix, daily trend, time-of-day distribution, movers, and AI crawls-vs-visits by company (default tab) |
/dashboard?view=bots |
Full bot list; add &bot=<name> for a per-bot detail view (trend chart, top pages, first/last seen) or &category=ai to filter to AI bots |
/dashboard?view=health |
2xx/3xx/4xx/5xx mix, failing paths, API/sensitive path hits |
/dashboard?view=events |
Filterable raw event log |
/api/bot-hit |
Authenticated bot event ingestion endpoint |
/login |
POST handler for token auth |
Older URLs (?view=ai, ?view=trends, ?view=status, ?view=pages, ?view=bot&bot=<name>) still work — they 307-redirect to their current equivalent (see src/proxy.ts).
Every view also accepts ?period=, either a preset (1, 7, 30, 90, 365 days) or a custom range as YYYY-MM-DD_YYYY-MM-DD. Periods over 90 days switch views into a rollup-backed "long-range mode" (see Architecture below) that hides path-level panels the daily rollup can't serve.
Tracked sites should POST JSON to the collector's /api/bot-hit endpoint with a project-scoped bearer token or x-bot-log-token header. Bot identity, category, and confidence are derived server-side from the submitted user agent and IP. The collector derives the project from the credential; a new credential cannot submit for another project.
Send status_code when it is available. Status reports show older or incomplete events as not captured; that is not a real HTTP status class.
curl -X POST "$DASHBOARD_URL/api/bot-hit" \
-H "Authorization: Bearer $BOT_INGEST_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"project": "marketing-site",
"environment": "production",
"url": "https://example.com/pricing?ref=ai",
"method": "GET",
"status_code": 200,
"user_agent": "GPTBot/1.0",
"ip": "203.0.113.10",
"referer": "",
"sample_rate": 1
}'Minimal non-blocking Next.js Proxy example:
import { NextResponse, type NextFetchEvent, type NextRequest } from "next/server";
import { isLikelyBotUserAgent } from "@/lib/bots";
function reportBotHit(request: NextRequest) {
return fetch(`${process.env.BOT_OBSERVABILITY_URL}/api/bot-hit`, {
method: "POST",
headers: {
"content-type": "application/json",
authorization: `Bearer ${process.env.BOT_INGEST_TOKEN}`,
},
body: JSON.stringify({
url: request.url,
method: request.method,
status_code: 0,
user_agent: request.headers.get("user-agent") ?? "",
ip: request.headers.get("x-forwarded-for")?.split(",")[0]?.trim() ?? "",
referer: request.headers.get("referer") ?? "",
}),
}).then((response) => {
if (!response.ok) console.error(`[bot-observability] ingestion failed: ${response.status}`);
});
}
export function proxy(request: NextRequest, event: NextFetchEvent) {
const response = NextResponse.next();
const userAgent = request.headers.get("user-agent") ?? "";
if (isLikelyBotUserAgent(userAgent)) {
event.waitUntil(reportBotHit(request).catch(() => undefined));
}
return response;
}Preserve each site's existing redirects, rewrites, locale handling, security headers, and content negotiation around this sender. The collector derives the project from the token, and status_code: 0 means that a pass-through Proxy does not claim to know the final downstream status.
Heartbeat events can be sent periodically to monitor pipeline freshness:
curl -X POST "$DASHBOARD_URL/api/bot-hit" \
-H "Authorization: Bearer $BOT_INGEST_TOKEN" \
-H "Content-Type: application/json" \
-d '{"heartbeat":true,"environment":"production"}'Heartbeats update one row in project_health per project. They are idempotent and do not append rows to bot_hits.
| Field | Required | Notes |
|---|---|---|
project or project_name |
Legacy only | New project-scoped credentials derive this value from the credential; legacy senders default to default |
url |
No | Used to derive host, path, and query_string when those are not provided |
path |
No | Useful when you do not want to send full URLs |
method |
No | Defaults to GET |
status_code or status |
No | Use the final HTTP response status when available |
user_agent |
Yes for bot detection | Non-bot events are ignored unless heartbeat is true |
ip |
No | Enables bot verification; stored only as a keyed HMAC-SHA-256 value, never as the raw submitted IP |
referer |
No | Stored for raw event inspection |
environment |
No | Defaults to production |
is_api_route |
No | Helps the Status tab surface API hits |
sample_rate |
No | Allowed values are 1, 0.5, 0.25, and 0.1; rollups weight each row by the exact integer reciprocal |
heartbeat |
No | Set true for pipeline health events |
The ingestion endpoint enforces a few hardcoded limits (see src/app/api/bot-hit/route.ts):
- Max request body: 32KB (
MAX_BODY_BYTES). Larger requests (byContent-Length) are rejected. - Rate limit: 120 requests/minute per caller IP (
RATE_LIMIT_RPM), tracked in an in-memory, per-serverless-instance store — see the caveat under Architecture; it is not a global ceiling on multi-instance platforms. - Field truncation: string fields are silently truncated, not rejected — 2000 characters for most string fields (
MAX_STRING_LENGTH), 1000 characters forpath(MAX_PATH_LENGTH). Oversized values are cut, not errored. - Responses:
Status Meaning 201Event stored ( { stored: true, bot_name, bot_category, confidence })200Not stored — non-bot, non-heartbeat traffic ( { stored: false, reason: "not_bot" })400Invalid or too-large JSON payload 401Missing/invalid ingestion credential 429Rate limit exceeded 503Ingestion not configured ( DATABASE_URL,BOT_IP_HASH_SECRET, or ingestion credentials missing/weak)
- The dashboard is protected by
BOT_ADMIN_TOKEN, not a full user-management system. A successful login creates a signed, HTTP-only, same-site session cookie valid for 1 year; the token itself is never stored in the cookie. - Ingestion uses project-scoped
BOT_INGEST_TOKENS; keep them server-side and never place them inNEXT_PUBLIC_*variables. - Submitted IP addresses are used for bot verification and then stored only as domain-separated, keyed HMAC-SHA-256 values derived from
BOT_IP_HASH_SECRET. Raw IP storage is not supported. - Rotating
BOT_IP_HASH_SECRETchanges the keyed hash produced for future observations of the same IP. Existing stored hashes remain unchanged. - User agents, paths, referrers, approximate geo fields, deployment URLs, and status codes may be stored.
- Rotate
DATABASE_URLand every role-specific secret before making a previously private deployment public if any may have been exposed outside trusted systems.
- Storage: a single
bot_hitsraw event table, two maintained tables —bot_hits_daily(a(day, project, bot, category, status_class)rollup used for long-range and high-volume views) andbot_first_seen(per-bot first/last-seen timestamps) — plus idempotentproject_healthheartbeat state. Bot events update the raw row and rollups in one transaction; heartbeats update onlyproject_health. If rows are ever ingested by an older build that predates the rollup,npm run reconcile-rollupsrebuilds both tables from raw (see Database Setup). Day buckets are UTC (DATE(created_at)); the UI displays timestamps in Europe/Berlin. - Request-scoped DB client: each request gets its own
postgresclient viacache()+after()(seesrc/lib/db.ts/src/app/dashboard/page.tsx), closed at the end of the request rather than pooled indefinitely — deliberate for small free-tier Postgres connection limits (e.g. Aiven). - Rendering: the dashboard streams server-rendered content with a
Suspenseboundary per view/panel, so slow queries don't block the whole page. - Rate limiting is per-instance, not global:
/api/bot-hit's rate limiter is an in-memoryMapscoped to a single running process (seesrc/app/api/bot-hit/route.ts). On multi-instance serverless platforms like Vercel, each concurrently-running instance enforces its own 120 req/min ceiling independently — there is no shared/global counter. Real aggregate throughput across all instances can therefore be significantly higher than 120 req/min. Do not rely on this limiter as a hard global cap; put a WAF/edge rate limit in front of it if you need one.
npm test/npm run test:unit— unit tests (pure logic: bot detection, category normalization, period parsing, attention-strip thresholds, etc.), no database required.npm run test:integration— integration tests against a real Postgres database, gated onTEST_DATABASE_URLbeing set (skipped otherwise). They apply the migrations, seed fixtures, and clean up after themselves.
Raw events can be retained for a bounded period while daily rollups, first/last-seen data, and project health remain available. The optional cleanup command defaults to 90 days:
npm run retain-raw
# or: node scripts/retain-raw-events.mjs "$DATABASE_URL" 90Schedule it daily or weekly. It only deletes old rows from bot_hits; it does not delete or rebuild rollups, first/last-seen records, or project_health:
DELETE FROM bot_hits
WHERE created_at < now() - interval '90 days';| Category | Description |
|---|---|
| AI Training | Bulk training data collectors (GPTBot, ClaudeBot, etc.) |
| AI Search | Indexers for AI chat products (OAI-SearchBot, PerplexityBot, etc.) |
| AI Agent | On-demand user-triggered fetches (ChatGPT-User, Claude-User, etc.) |
| Search Engine | Traditional search index crawlers (Googlebot, Bingbot, etc.) |
| Social Preview | Link unfurling / share preview cards (Twitterbot, Slack, etc.) |
| SEO Tool | SEO audit and tech detection (Ahrefs, Semrush, etc.) |
| Monitoring | Uptime / performance checks (Pingdom, UptimeRobot, etc.) |
| Archival | Web page preservation (Internet Archive, etc.) |
| Generic / CLI | Uncategorized automated agents (curl, wget, etc.) |