Collecting data

An artifact can store what a visitor types, in a database that comes with it. One field on the deploy switches it on.

An artifact is a static bundle, and that used to be the end of the sentence. It is not: a page published here can collect what a visitor types into it — a survey answer, an RSVP, a tracker entry, an order — and the agent that built the page reads the rows back with its own key. The database comes with the artifact. There is nothing to provision, host or sign up for.

It is off until the deploy declares a policy. That is the one thing worth knowing before anything else, because an artifact without a policy renders its form, runs its submit handler, and drops every answer — and nothing about the page looks wrong while it does. The first sign is an empty read a week later.

What to ask your agent for

Ask for the thing you want. "A page to collect RSVPs for my son's birthday — kid's name, parent's name, and whether they're coming." "A sign-up sheet for the Saturday shift." "Ask the team how the retro went and show me the answers."

Your agent does not need to be told about policies, tokens or endpoints. All of it is in the skill and the MCP server's tools.

If an agent tells you this platform cannot store form responses, it is wrong and its copy of the skill is stale. Ask it to re-fetch SKILL.md from the platform that issued its key and read the Collecting what a visitor types section. The answer is never a mailto: link, a WhatsApp message, or a Google Form — each of those hands your data to somebody else and loses the read-back that is the point of collecting it here.

What a policy declares

Four decisions and a list of fields, sent as policy alongside metadata and files on the deploy. The deploy contract has the wire format; what the decisions mean:

  • read — who may read a row back from inside the page. admin is nobody but you, which is right for an RSVP list. own shows a visitor only their own rows, which is right for a tracker. contributors and public show them everybody's, and public in particular is a decision to make rather than a default to inherit. None of these limit you: your key reads every row regardless.
  • perUser — one row per person, or many. single is one row per person; many is one row per thing a person did. Which you want is decided by the shape of the data, below.
  • writes — whether a visitor may change a row that exists. Follows from the shape: upsert for an answer sheet or a person's own editable records, append-only for an event log. It is not a free toggle. Erasure works either way.
  • retentionDays — how long a row is kept. A finite number is required for anything classified sensitive, and expired rows are pruned rather than merely hidden.
  • fields — the shape of one submission, each with a type and a pii classification: none, contact (an email, a handle, a phone number) or sensitive. pii has no default, on purpose. A field's type may never change between versions.

The deploy response says what was actually accepted, including any field the platform reclassified on its own — an unconstrained free-text box declared pii: none is stored and masked as contact whatever the declaration said, because that is the field most likely to end up holding an email address.

Shape the data before you ask for it

Columns are questions; rows are people. Ask one thing first: is the unit a person or an event?

| The unit is | Shape | You read it as | | --- | --- | --- | | a person — a form, a quiz, onboarding, a checklist, four decisions on one page | answer sheet: one row per person, one column per question, filled in as they go (perUser: single, writes: upsert) | a spreadsheet | | a thing that happened — a guestbook, an order, a workout set, a vote | event log: one complete row per event (perUser: many, writes: append-only) | a timeline | | an item a person keeps and edits — a todo, a journal entry | editable records: the author can replace or delete a row by its id (perUser: many, writes: upsert) | a list |

The mistake worth knowing about is one row per tap. A page that asks four questions and stores each answer as a row with a question column and an answer column hands you five rows from two people, with the questions as values. Declare one optional field per question instead: each step saves only the field it answers, the first save creates the row, and every later one merges into it. A known set of questions is columns; open-ended repeated items are rows. A policy that looks like the first mistake still publishes, with a warning your agent reads and can fix. Agents are taught this shape by the skill, so you should not have to ask for it.

Reading it back

Three ways, all reading the same rows:

  • The dashboard, at /data. Every artifact of yours that collects, what it has, and an export.
  • Your agent, over MCP — super_data with artifact for one artifact, or with collection for a group that shares a collection label, read over one window.
  • A visitor, from inside the page, seeing only what the read policy allows.

Prefilling a form from the visitor's own answers

A form that says "resubmitting updates your answers" has to show them, or a second submit overwrites the first with the form's defaults. From inside the page, GET the responses address with the page's data token and ?scope=own:

const { token, artifact } = await fetch(location.pathname.replace(/\/$/, '') + '/__sa/data-token')
  .then((r) => r.json())
const url = '/api/data/artifacts/' + artifact + '/responses'
const auth = { authorization: 'Bearer ' + token }

const { responses } = await fetch(url + '?scope=own&limit=1', { headers: auth })
  .then((r) => r.json())
const mine = responses[0]?.payload ?? {}   // {} for a first-time visitor
form.attending.value = mine.attending ?? ''

// on submit: 201 the first time, 200 every time after, both mean saved
const res = await fetch(url, {
  method: 'POST',
  headers: { 'content-type': 'application/json', ...auth },
  body: JSON.stringify({ attending: form.attending.value }),
})
status.textContent = res.ok ? 'Your answers were saved.' : 'Not saved. Try again.'
  • Only their own rows. The token names the visitor, and no parameter can name another. Everyone else's answers stay out of reach whatever read says.
  • The policy must allow it. read: "own" (or contributors, or public) lets a visitor read; the default admin answers 403, so a form meant to prefill declares own.
  • Who counts as the same visitor. A signed-in person sees their answers on any device. An anonymous visitor sees them in the same browser; a new browser, or cleared cookies, is a new visitor with nothing to prefill.

A visitor erases their own rows themselves. Deleting the artifact closes the data with it: the tombstone is written before a single byte is removed, so a write token minted seconds earlier cannot land afterwards.

What this is not

It is not analytics. Analytics is what the platform observes about a request — views, devices, errors. Collected data is what the page asked for and the visitor chose to give. They are stored, read and erased separately, and neither answers the other's question.

It is not a database you get to query freely. One submission is one flat row of declared fields, capped in size and in count. That constraint is what makes the rows legible to the next version, to the dashboard and to an agent that did not write them.

Collecting data — Super Artifacts docs