Corpusforge
A local-first studio where an AI agent plans, writes and verifies conversational training data grounded in real sources, then hands you a clean dataset in the format your training stack reads.

8/8
conversations accepted on the recorded run
84.7
average quality score out of 100
3:11
minutes for the whole job on a free-tier quota
$0.02
total cost of 70 grounded turns
Inside the product
See it working
Real screens and a recorded run, captured from the working product.
Studio home: "Turn any source into training data you can trust", with the animated sources-forge-verified scene.
01 / 16My role
Sole engineer: the planning agent, the generation and verification pipeline, the rate limiter, the live monitor and the bilingual interface.
The problem
Fine-tuned models are only as good as their worst training conversation, and raw model output drifts: it invents facts, leaks phrases like "according to the text" and repeats itself. Writing conversations by hand does not scale, and a single prompt gives no way to check or repair what comes out.
What I built
Corpusforge treats generation as a pipeline with a quality gate. A tool-using agent turns a plain-language goal into a concrete job (topics, conversation count, turn range, answer length, personas, styles, cost and time estimate) and nothing runs until you press Launch.
Conversations are written from material, not memory: your own files, web pages imported by URL and stripped of menus, ads and footers, pages the agent researched, or model knowledge for low-risk topics. Every turn passes deterministic checks (structure, source privacy, language and script, variety, answer length, style and grounding); failures go back to the model as repair feedback, and each conversation is scored from 0 to 100 and lands as accepted, flagged or rejected under thresholds you control. You review the evidence with one-key Accept, Flag or Reject and export to OpenAI JSONL, ShareGPT, Alpaca, CSV, Excel or Parquet.
Architecture
Python FastAPI backend with async SQLAlchemy 2 on SQLite (WAL), server-sent events for the live monitor, trafilatura for clean web extraction, PyMuPDF and python-docx for files, PyArrow for Parquet, and semantic search over sources with gemini-embedding-2. A Next.js 16 static export with React 19, Tailwind CSS and Motion is served by the API. Works with Gemini out of the box or any OpenAI-compatible endpoint (OpenAI, Ollama, LM Studio, vLLM, llama.cpp). One cross-platform task script sets up, starts, checks and smoke-tests everything.
Challenges
Running long jobs on strict free-tier limits without retry storms. An adaptive rate limiter combines a sliding request and token window, an AIMD concurrency gate and a global cool-down on 429 retry hints, so a quota burst becomes a short pause. Jobs pause, resume and survive restarts. Generation and quality checks also had to handle Arabic properly (normalisation, script and diacritics), while conversations can be written in dialects such as Jordanian Arabic.
Key engineering decisions
Deterministic checks by default instead of a model-as-judge, so quality is cheap, explainable and repeatable. The agent proposes but never launches on its own. Every accepted answer keeps the source passages behind it, so any figure in the dataset can be traced.
Results & impact
In the recorded run the agent read two sources (an English solar guide and an Arabic page imported by URL and cleaned of its cookie banner, navigation, ad and footer) and planned 8 customer-support conversations in English and Jordanian Arabic for $0.0005. The job finished in 3 min 11 s at 15 requests per minute: 8 of 8 accepted, 70 turns, average quality 84.7 (78.6 to 88.7), 53 requests, 114.9K tokens and $0.02. The repair loop rewrote turns the grounding check found unsupported before acceptance. A second job held 4 weaker conversations back for review (scores 64 to 66) instead of passing them, and paused for 42 s when the provider asked instead of failing.
Highlights
- Planning agent with tools: suggest topics, research the web, read and search sources, review the dataset, propose a job
- Three grounding modes: research, your own files and URLs, or topics only
- Clean ingestion that strips navigation, ads, cookie banners and footers
- Quality gate with a repair loop, 0 to 100 scoring, thresholds and deduplication
- Live job monitor over server-sent events with an animated sources-to-forge-to-verified scene
- Review drawer with per-dimension quality, evidence passages and keyboard review
- Export to OpenAI chat JSONL, ShareGPT, Alpaca, CSV, Excel and Parquet
- English and Arabic UI with RTL, dark and light themes, command palette and phone layout

