AI Metadata Generation — SCQH
AI Metadata Generation — SCQH
Situation
About a year ago we shipped AI metadata generation in PortalJS Cloud: when a publisher uploads files, it drafts the dataset's metadata for them, and we regularly show it off when selling projects.
Complication
It hasn't been touched since launch, and there's clear room to make it better — it runs on an outdated model, the metadata it produces is often unreliable, and the publishing UX is basic — so it's a weaker demo and tool than it could be.
Question
How do we improve its quality and UX enough to make it a compelling demo when we're selling projects — and make the same capability usable beyond the browser?
Hypothesis
Refreshing the model, grounding inference in the actual uploaded data so the output is accurate, and polishing the publish-flow UX will make it a noticeably better, more trustworthy demo — and exposing the same engine through an API and a reusable skill makes it useful outside the GUI too.
Issue tree — breaking down the Question
The Question splits into four things we need to work out:
-
Why does the metadata look unreliable in a demo?
- Slugs and names aren't validated against CKAN's rules, so saving can fail outright in front of a prospect.
- Inference runs on only ~2 KB of each file, so values like coverage, dates, and category lists are visibly wrong on anything but a tiny file.
- Useful fields like tags, groups, and descriptions come back empty, so the form still looks half-finished.
- The model is old.
-
What's weak about the UX?
- A single static spinner with no progress during generation, which feels broken on larger files.
- No signal of which fields the AI was confident about vs. guessing, so a reviewer can't tell what to trust.
-
Which interfaces should it serve?
- GUI. The browser flow is what we demo; it has to look good.
- API. Many customers publish programmatically and in bulk, not one dataset at a time.
- Skill. Packaged as a skill (e.g. usable from Claude Code), the same capability works in agent-driven workflows.
- Live in-flow agent. A conversational assistant that refines metadata interactively — an idea worth noting.
-
What should power it?
- Which model/provider to move to, and whether to abstract the provider so we aren't locked in.
- Where heavy file parsing runs (browser vs. server), which also determines how sensitive data is handled.
Hypothesis tree — sub-hypotheses to test
- H1 — Correctness floor (cheap, deterministic). Validating and auto-repairing values (slug characters, unique names) before saving removes the errors that can break a live demo.
- H2 — Data-grounded inference. Computing real statistics over the full file (e.g. DuckDB-WASM SQL in the browser, or server-side aggregates) and feeding them to the model makes coverage, dates, and descriptions accurate — and, done in-browser, keeps customer data local.
- H3 — Richer auto-population. Grounding the model in the portal's existing tags and groups lets it fill those fields, so the generated metadata looks complete.
- H4 — Model/provider refresh. Moving off the old model behind a provider abstraction improves quality and opens options, including privacy-preserving ones.
- H5 — UX polish. Clear progress feedback and per-field confidence signals make the flow feel solid and make "review before publishing" meaningful.
- H6 — Reusable engine (API + skill). Factoring the generation logic into a shared core that the GUI, an API, and a skill all call lets the same capability serve programmatic and agent-driven publishing — and avoids duplicating the parallel AI-agent framework work.
Idea, not a sub-hypothesis: a live in-flow agent that lets a publisher refine metadata conversationally — noting it for discussion.