Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
b67139c
add provider choice (or concepts of it at least)
AdamEXu Jun 9, 2026
8cf6bea
tested a little bit and it should work i think? idk
AdamEXu Jun 9, 2026
ef9de13
testing row extraction
AdamEXu Jun 11, 2026
56f68ab
Address PR review: hard-validate model slugs + auth-guard provider mo…
AdamEXu Jul 15, 2026
3839441
Fix model-save false-rejections found in self-review
AdamEXu Jul 15, 2026
e20eb53
Merge branch 'main' into model-choice
AdamEXu Jul 22, 2026
241e401
Modernize default model slugs + fix provider-selector reset
AdamEXu Jul 22, 2026
d1ae1ef
Address provider configuration review findings
AdamEXu Jul 22, 2026
eb44dbf
Add OpenRouter app attribution headers
AdamEXu Jul 22, 2026
eeb8116
Refresh default model slugs to current provider lineups
AdamEXu Aug 4, 2026
7d560fc
Re-verify every default model against provider docs
AdamEXu Aug 4, 2026
4b2b169
Add per-role reasoning effort control
AdamEXu Aug 4, 2026
51961d9
Add reasoning effort to setup wizard
AdamEXu Aug 4, 2026
5c25efb
Merge branch 'main' into model-choice
AdamEXu Aug 6, 2026
14d8b69
Remove playwright row-extractor spike from main
AdamEXu Aug 6, 2026
502aaad
Merge remote-tracking branch 'upstream/main'
AdamEXu Aug 6, 2026
f992352
Update dependencies to clear security advisories
AdamEXu Aug 6, 2026
5b0f589
Ship frontend/lib in the release package and verify it
AdamEXu Aug 6, 2026
debcd9f
Fix Fireworks AI API key link
AdamEXu Aug 6, 2026
06f0cac
Replace the reasoning slider with a segmented control
AdamEXu Aug 6, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 5 additions & 5 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,9 +6,9 @@ Frontend on :3500, backend on :3501, Mastra Studio on :4111, Convex dashboard on

## Setup

1. Copy `.env.example` to `.env` and fill in your keys:
1. Local mode creates `.env` automatically and collects TinyFish + LLM provider credentials in the setup UI. Production still uses env keys:
- `TINYFISH_API_KEY` — from https://agent.tinyfish.ai/api-keys?utm_source=github&utm_medium=organic&utm_campaign=bigset-developer-2026q2
- `OPENROUTER_API_KEY` — from https://openrouter.ai/settings/keys
- `OPENROUTER_API_KEY` — production default LLM provider key
- `NEXT_PUBLIC_CLERK_PUBLISHABLE_KEY` — from Clerk API Keys
- `CLERK_SECRET_KEY` — from Clerk API Keys
- `CLERK_JWT_ISSUER_DOMAIN` — your Frontend API URL (e.g. `https://your-app.clerk.accounts.dev`)
Expand All @@ -26,17 +26,17 @@ Frontend uses Convex React hooks (`useQuery`, `useMutation`) with `ConvexProvide

Backend is Fastify + Mastra. Fastify serves the HTTP API (Clerk JWT auth on protected routes via `backend/src/clerk-auth.ts`). Mastra (`backend/src/mastra/`) is the workflow orchestration layer — it wraps pipelines into inspectable workflows with a Studio UI. Both run as separate Docker services sharing the same source code.

The schema inference pipeline: frontend calls `POST /infer-schema` → Fastify verifies the Clerk JWT → calls `inferSchema()` in `backend/src/pipeline/schema-inference.ts` → Claude Sonnet 4.6 via OpenRouter → returns a Zod-validated `DatasetSchema` → frontend maps it to editable columns in the wizard.
The schema inference pipeline: frontend calls `POST /infer-schema` → Fastify verifies the Clerk JWT → calls `inferSchema()` in `backend/src/pipeline/schema-inference.ts` → selected LLM provider/model → returns a Zod-validated `DatasetSchema` → frontend maps it to editable columns in the wizard.

The populate pipeline: frontend calls `POST /populate` with `{ datasetId, datasetName, description, columns }` → Fastify verifies the Clerk JWT → triggers `populateWorkflow` which: (1) clears existing rows, (2) builds a prompt from the schema, (3) runs the populate agent (Claude Sonnet 4.6) which searches the web via TinyFish APIs, then inserts rows into Convex one by one. Rows appear in realtime on the frontend via Convex reactive queries.
The populate pipeline: frontend calls `POST /populate` with `{ datasetId, datasetName, description, columns }` → Fastify verifies the Clerk JWT → triggers `populateWorkflow` which: (1) clears existing rows, (2) builds a prompt from the schema, (3) runs the populate agent using the selected LLM provider/model. The agent searches the web via TinyFish APIs, then inserts rows into Convex one by one. Rows appear in realtime on the frontend via Convex reactive queries.

Convex functions use `ctx.auth.getUserIdentity()` to get the authenticated user. The `ownerId` field on datasets stores `identity.subject` (Clerk user ID). Do not pass `ownerId` from the client.

## Environment Variables

Root `.env` is the only local env file. Docker Compose, package scripts, and Convex CLI helper targets all read it. Key variables:
- `TINYFISH_API_KEY` — used by the populate agent for web search and fetch (get one at https://agent.tinyfish.ai/api-keys?utm_source=github&utm_medium=organic&utm_campaign=bigset-developer-2026q2)
- `OPENROUTER_API_KEY` — used by backend and Mastra for AI model calls
- `OPENROUTER_API_KEY` — production default LLM provider key; local mode can use OpenRouter, OpenAI, Anthropic, or custom OpenAI-compatible via setup UI
- `NEXT_PUBLIC_CLERK_PUBLISHABLE_KEY`, `CLERK_SECRET_KEY` — shared by frontend and backend
- `CONVEX_SELF_HOSTED_ADMIN_KEY` — used by backend for system-level Convex writes

Expand Down
27 changes: 14 additions & 13 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -93,7 +93,7 @@ On first launch, BigSet sends you to setup. You'll connect two services:
| Service | What it's for | Get your key |
|---------|--------------|-------------|
| **TinyFish** | Web search + page fetching | [tinyfish.ai/api-keys](https://agent.tinyfish.ai/api-keys?utm_source=github&utm_medium=organic&utm_campaign=bigset-developer-2026q2) |
| **OpenRouter** | LLM calls (schema inference + agents) | [openrouter.ai/settings/keys](https://openrouter.ai/settings/keys) |
| **LLM provider** | Schema inference + agents | OpenRouter, direct providers, Ollama, LM Studio, or another OpenAI-compatible endpoint |

Local API keys are stored in your OS keychain.

Expand Down Expand Up @@ -214,23 +214,24 @@ Once everything is ready, you'll see:
| **Mastra Studio** (workflow inspector) | [localhost:4111](http://localhost:4111) |

Open [localhost:3500](http://localhost:3500). The setup screen will ask for
TinyFish and OpenRouter credentials and save them to your OS keychain for this
workspace.
TinyFish credentials plus an LLM provider. Choose OpenRouter, a direct hosted
provider, Ollama, LM Studio, or another OpenAI-compatible endpoint. BigSet saves
local keys to your OS keychain for this workspace.

### Step 3: Connect TinyFish and OpenRouter
### Step 3: Connect TinyFish and an LLM provider

TinyFish powers web search and page fetching. OpenRouter routes LLM calls to
the models BigSet uses for schema inference and agents.
TinyFish powers web search and page fetching. Your LLM provider powers schema
inference and dataset-building agents.

1. Create a TinyFish key at [agent.tinyfish.ai/api-keys](https://agent.tinyfish.ai/api-keys?utm_source=github&utm_medium=organic&utm_campaign=bigset-developer-2026q2)
2. Create an OpenRouter key at [openrouter.ai/settings/keys](https://openrouter.ai/settings/keys)
3. Paste both into BigSet's setup screen

OpenRouter is pay-as-you-go; $5-10 is plenty to start.
2. Choose an LLM provider. Hosted providers use an API key; Ollama and LM Studio
need no key by default.
3. Save the provider, then choose models for each role in the model picker when
the provider does not supply defaults.

> **Note:** root `.env` is the only local env file. If you edit Convex functions in `frontend/convex/`, run `make convex-push` to deploy the changes.

> **Free tier:** cloud signed-in accounts get **2,500 row operations per calendar month** (resets on the 1st, UTC). Local mode bypasses the cloud quota and uses your TinyFish/OpenRouter accounts directly.
>
> **Free tier:** cloud signed-in accounts get **2,500 row operations per calendar month** (resets on the 1st, UTC). Local mode bypasses the cloud quota and uses your TinyFish + LLM provider accounts directly.

### Step 4 (optional): Load curated datasets

Expand Down Expand Up @@ -311,7 +312,7 @@ If you want a completely fresh start: `make clean` then `make dev`.
| Auth | Local auth (dev); [Clerk](https://clerk.com) (cloud) |
| Database | [Convex](https://convex.dev) (self-hosted) |
| Data Collection | [TinyFish](https://www.tinyfish.ai?utm_source=github&utm_medium=organic&utm_campaign=bigset-developer-2026q2) APIs (Search, Fetch, Browser) |
| AI orchestration | [Mastra](https://mastra.ai) workflows + [Vercel AI SDK](https://sdk.vercel.ai) + [OpenRouter](https://openrouter.ai) → Claude Sonnet (schema inference + populate agent) |
| AI orchestration | [Mastra](https://mastra.ai) workflows + [Vercel AI SDK](https://sdk.vercel.ai) + a configured hosted, local, or OpenAI-compatible LLM provider |
| Table view | [TanStack Table](https://tanstack.com/table) + [react-window](https://github.com/bvaughn/react-window) virtualization |
| Exports | CSV (built-in) + XLSX ([SheetJS](https://sheetjs.com), dynamic-imported) |
| Analytics | [PostHog](https://posthog.com) — events, session replay, error tracking (optional) |
Expand Down
6 changes: 3 additions & 3 deletions backend/CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ To add a new protected route, register it inside the scoped plugin in `src/index

## Schema Inference Pipeline

`src/pipeline/schema-inference.ts` — takes a natural language prompt and returns a structured `DatasetSchema` (Zod-validated, defined in `src/pipeline/types.ts`). Uses Claude Sonnet 4.6 via OpenRouter (`@openrouter/ai-sdk-provider` + Vercel AI SDK). Auto-retries once on validation failure by feeding the error back to the model.
`src/pipeline/schema-inference.ts` — takes a natural language prompt and returns a structured `DatasetSchema` (Zod-validated, defined in `src/pipeline/types.ts`). Uses the configured LLM provider/model via `src/config/llm.ts` + Vercel AI SDK. Auto-retries once on validation failure by feeding the error back to the model.

The pipeline is a pure function (`inferSchema(prompt) → DatasetSchema`). It is called by both Fastify (for the HTTP API) and Mastra (for workflow orchestration).

Expand All @@ -26,7 +26,7 @@ The pipeline is a pure function (`inferSchema(prompt) → DatasetSchema`). It is
- `src/mastra/index.ts` — registers workflows with the `Mastra` instance (the populate agent is built per-run, not registered as a singleton)
- `src/mastra/workflows/infer-schema.ts` — `inferSchemaWorkflow`, a single-step workflow wrapping `inferSchema()`
- `src/mastra/workflows/populate.ts` — `populateWorkflow`, 3-step workflow: clear rows → build prompt → run populate agent
- `src/mastra/agents/populate.ts` — `buildPopulateAgent(authorizedDatasetId, authContext, columns)`, builds the orchestrator agent (Claude Sonnet 4.6) with 3 tools: `search_web`, `fetch_page`, `investigate_row`. No write access — all inserts go through investigate subagents.
- `src/mastra/agents/populate.ts` — `buildPopulateAgent(authorizedDatasetId, authContext, columns)`, builds the orchestrator agent with the configured LLM provider/model and 3 tools: `search_web`, `fetch_page`, `investigate_row`. No write access — all inserts go through investigate subagents.
- `src/mastra/agents/investigate.ts` — `buildInvestigateAgent(authorizedDatasetId, authContext, columns)`, builds a per-entity subagent with `insert_row`, `list_rows`, `search_web`, `fetch_page`. Researches one entity, inserts at most one row, returns structured feedback (`INSERTED/SUMMARY/CLUES/REASON`).
- `src/mastra/tools/investigate-tool.ts` — `buildInvestigateTool(authorizedDatasetId, authContext, columns)` creates the `investigate_row` tool. The orchestrator calls it to hand off a lead; it spawns a fresh investigate agent, runs it (maxSteps: 25), parses the structured output, and returns it to the orchestrator. Errors are caught and returned as structured failures so the orchestrator can self-correct.
- `src/mastra/tools/dataset-tools.ts` — `buildPopulateTools(authorizedDatasetId, authContext)` factory returning 5 Convex-backed tools: `insert_row`, `list_rows`, `get_row`, `update_row`, `delete_row`. The dataset id is captured by closure so the LLM cannot redirect writes to other datasets; `authContext` (Clerk userId + workflow run id) is captured for caller-attribution in security logs and the `CAPABILITY_VIOLATION` PostHog event. See the security note at the top of the file.
Expand All @@ -48,7 +48,7 @@ Required env vars (see `.env.example`):
- `CONVEX_URL` — Convex instance URL
- `CONVEX_SELF_HOSTED_ADMIN_KEY` — for system-level Convex writes (internal mutations)
- `CLERK_SECRET_KEY`, `CLERK_PUBLISHABLE_KEY` — for JWT verification
- `OPENROUTER_API_KEY` — for AI model calls
- `OPENROUTER_API_KEY` — production default LLM provider key; local mode can use providers including OpenRouter, direct hosted APIs, Ollama, LM Studio, and custom OpenAI-compatible endpoints via the setup UI
- `TINYFISH_API_KEY` — for web search and fetch (populate agent). Get one at https://agent.tinyfish.ai/api-keys?utm_source=github&utm_medium=organic&utm_campaign=bigset-developer-2026q2

In Docker, these are interpolated from the root `.env` file via `docker-compose.dev.yml`.
Loading