CrawlForge
A careers page in. Structured job intelligence out.
Live · 20 tests · Updated 2026
Java 21 · Spring Boot · Spring JDBC · H2 · jsoup
Hosted on Render. The first open may take up to about 60 seconds; the proof below remains available immediately.
- Context
- Independent product project · 2026
- My role
- Product designer and sole engineer
- Team
- Solo build
- Evidence
- 20 tests · persistent bounded BFS
01 Problem
Company careers sites and ATS pages expose the same job data through inconsistent HTML, embedded JSON-LD, redirects, and deep link graphs. A useful crawler has to find the listings, normalize them, and return evidence a person or downstream system can actually use.
02 Constraints
- 01Private and local network targets are blocked before requests and after redirects.
- 02Crawls are bounded by page count, depth, host scope, robots.txt, and per-host request rate.
- 03The breadth-first frontier and visited URLs persist so interrupted runs can recover.
- 04Job extraction prefers JobPosting JSON-LD and falls back to conservative HTML signals.
- 05Exports and deterministic resume matching remain available without an AI provider key.
03 Architecture
04 Decisions
Discovery and extraction are separate stages.
Why The crawler can reason about URLs and recovery without coupling every fetch to one careers-site layout.
Tradeoff There are more persisted states, but parser changes do not rewrite the crawl engine.
Structured data wins over page chrome.
Why JobPosting JSON-LD carries title, location, description, and hiring metadata with less navigation noise than rendered text.
Tradeoff Malformed or missing structured data still needs a conservative HTML fallback.
AI is optional; scoring is not.
Why A live demo and API integration must still return repeatable matches when no provider key is configured.
Tradeoff The deterministic score is less nuanced, while the optional model is limited to grounded explanation.
05 Failure
A target resolves to a private address.
Validation rejects it before retrieval; redirects are checked again.
The careers page links into an infinite graph.
Canonicalization, deduplication, depth, and page limits bound the crawl.
The process restarts mid-run.
The frontier and visited state remain in H2 so work can resume.
The page has navigation text but no job schema.
The extractor avoids inventing a job and records parser diagnostics.
AI matching is not configured.
Deterministic score, matched skills, gaps, and exports continue to work.
06 Object
live ZoomInfo scan · one structured posting
{
"title": "Senior Software Engineer, Data Infrastructure",
"location": "US; Remote",
"experience": "5+ years of professional software engineering experience",
"skills": ["Java", "Python", "AWS", "Kubernetes", "Kafka", "Distributed Systems"]
}
07 Inspect
08 What I would improve next
- Add a browser-rendering adapter for JavaScript-only careers sites without weakening the default HTTP crawler.
- Move crawl state to PostgreSQL and use database leases or fencing tokens for safe multi-worker ownership.
- Expose crawl budgets, parser coverage, host latency, and retry causes as operational metrics.