A URL is the beginning, not the input format
A company careers page may contain job cards, links to an applicant-tracking system, embedded JobPosting JSON-LD, or almost no useful HTML at all. It may redirect across domains, repeat the same listing under tracking parameters, or link into an effectively unbounded corporate site. Treating that page as a single scrape works only in the happy-path demo.
CrawlForge starts from a different contract: accept one public Careers URL, discover a bounded set of candidate pages, extract only supported evidence, and return structured jobs plus diagnostics. Discovery and extraction are deliberately separate. One decides which URLs deserve attention; the other decides whether a retrieved document contains a defensible job posting.
The crawler's product guarantee is not “fetch everything.” It is “do useful work inside explicit limits.”
Why breadth-first search fits careers discovery
Careers sites tend to place useful pages close to the seed: a jobs index, department pages, pagination, then job detail pages. Breadth-first search visits that neighborhood before spending the budget on deep navigation chains. That makes the result easier to reason about than an unconstrained recursive walk.
The frontier is persisted rather than held only in memory. Each URL carries depth and lifecycle state. If the process stops after a fetch but before parsing completes, the database still contains the work and its last known state. Recovery becomes a state transition, not a fresh crawl that silently duplicates requests.
PENDING → FETCHING → DONE
↘ RETRY
↘ SKIPPED
↘ FAILED
Bounds are part of the API
Maximum pages and maximum depth are visible inputs, not hidden implementation constants. Host scope prevents a careers crawl from wandering into unrelated domains. Canonicalization removes fragments and known tracking noise before deduplication, so two marketing URLs do not consume two units of crawl budget.
Network safety has to survive redirects. Blocking a private address only before the first request is insufficient if a public URL can redirect to localhost or an internal range. CrawlForge validates the initial target and every redirect, caps redirect count, limits response size, applies per-host rate limits, and can respect robots.txt. These controls serve different purposes: robots expresses site policy; SSRF checks protect the service; budgets protect both sides from accidental excess.
- Page and depth ceilings make completion predictable.
- Canonical URLs and a visited set make duplicate work visible and avoidable.
- Host boundaries, redirect validation, and response limits keep the crawler inside its authority.
- A bounded retry policy handles transient failure without creating an immortal request.
Extract evidence before inference
Once a page is retrieved, the extractor prefers JobPosting JSON-LD. Structured data usually provides a cleaner title, location, description, employment type, and hiring organization than a page-wide text scrape. The HTML fallback is intentionally conservative. Navigation copy that happens to contain words such as engineering or remote is not enough to invent a job.
Skills and experience are normalized from the evidence that was actually found. JSON and CSV exports remain available without an AI key. Resume matching also has a deterministic base score, matched skills, and gaps. An optional model may explain that evidence, but it cannot be the only path that makes the product work.
Test the graph, not just the parser
Parser unit tests are necessary but they miss the system behavior that makes this a crawler. CrawlForge's 20-test suite includes a local end-to-end link graph, tracking-parameter canonicalization, robots exclusion, bounded traversal, and a transient 503 that succeeds on retry. Those cases test ownership of the frontier and the conditions under which a run stops.
A useful failure report should answer: Was the URL rejected for safety? Skipped by policy? Retried after a transient response? Parsed without a supported job? A total count without those distinctions turns operations into guesswork.
What I would improve next
The current HTTP-first crawler intentionally does not pretend to render every JavaScript-only job board. I would add a browser adapter behind the same retrieval boundary, enabled only for supported providers or explicit cases. The default path should stay cheap, observable, and easy to test.
For multiple workers, I would move H2 state to PostgreSQL and add database leases or fencing tokens so only the current owner can complete frontier work. I would also publish parser coverage, host latency, retry causes, and remaining crawl budget as first-class metrics. The next version should not merely crawl more sites; it should make every additional capability easier to operate.