← Selected work CrawlForge / Case study Source ↗

CrawlForge

A careers page in. Structured job intelligence out.

Live · 20 tests · Updated 2026

Java 21 · Spring Boot · Spring JDBC · H2 · jsoup

Hosted on Render. The first open may take up to about 60 seconds; the proof below remains available immediately.

Context
Independent product project · 2026
My role
Product designer and sole engineer
Team
Solo build
Evidence
20 tests · persistent bounded BFS

01 Problem

Company careers sites and ATS pages expose the same job data through inconsistent HTML, embedded JSON-LD, redirects, and deep link graphs. A useful crawler has to find the listings, normalize them, and return evidence a person or downstream system can actually use.

02 Constraints

  1. 01Private and local network targets are blocked before requests and after redirects.
  2. 02Crawls are bounded by page count, depth, host scope, robots.txt, and per-host request rate.
  3. 03The breadth-first frontier and visited URLs persist so interrupted runs can recover.
  4. 04Job extraction prefers JobPosting JSON-LD and falls back to conservative HTML signals.
  5. 05Exports and deterministic resume matching remain available without an AI provider key.

03 Architecture

Careers intelligence path

04 Decisions

Discovery and extraction are separate stages.

Why The crawler can reason about URLs and recovery without coupling every fetch to one careers-site layout.

Tradeoff There are more persisted states, but parser changes do not rewrite the crawl engine.

Structured data wins over page chrome.

Why JobPosting JSON-LD carries title, location, description, and hiring metadata with less navigation noise than rendered text.

Tradeoff Malformed or missing structured data still needs a conservative HTML fallback.

AI is optional; scoring is not.

Why A live demo and API integration must still return repeatable matches when no provider key is configured.

Tradeoff The deterministic score is less nuanced, while the optional model is limited to grounded explanation.

05 Failure

  • A target resolves to a private address.

    Validation rejects it before retrieval; redirects are checked again.

  • The careers page links into an infinite graph.

    Canonicalization, deduplication, depth, and page limits bound the crawl.

  • The process restarts mid-run.

    The frontier and visited state remain in H2 so work can resume.

  • The page has navigation text but no job schema.

    The extractor avoids inventing a job and records parser diagnostics.

  • AI matching is not configured.

    Deterministic score, matched skills, gaps, and exports continue to work.

06 Object

live ZoomInfo scan · one structured posting

{ "title": "Senior Software Engineer, Data Infrastructure", "location": "US; Remote", "experience": "5+ years of professional software engineering experience", "skills": ["Java", "Python", "AWS", "Kubernetes", "Kafka", "Distributed Systems"] }

07 Inspect

Live Open product ↗

Source GitHub ↗

Wake time Allow up to about 60 seconds on the first request.

08 What I would improve next

  1. Add a browser-rendering adapter for JavaScript-only careers sites without weakening the default HTTP crawler.
  2. Move crawl state to PostgreSQL and use database leases or fencing tokens for safe multi-worker ownership.
  3. Expose crawl budgets, parser coverage, host latency, and retry causes as operational metrics.