Gil Allouche
← All writing
TutorialEntrepreneurship in the age of AI

Build a SaaS app with AI, starting with the token meter

A working multi-tenant API with key auth, per-tenant token metering, a hard quota and one AI endpoint — about 30 minutes and ~120 lines of Node.

September 29, 2026·Gil Allouche·13 min read
A build-along. Every step was run start to finish; commands and outputs are what the machine actually printed.

The whole build is 120 lines of Node and one SQLite file, and the interesting part is four numbers: $0.15 per 1M input tokens, $0.60 per 1M output tokens (OpenAI pricing), a 100,000-token monthly cap, and HTTP 402 — the status code reserved for "you owe money" since RFC 9110 was published in June 2022. Output costs 4x input on gpt-4o-mini, which is why the usage table below stores the two counts in separate columns and never sums them before writing, and why the enforcement check runs before the model call instead of after it.

What you end up with

A running HTTP API on localhost:8787 that accepts a support ticket, returns a two-sentence summary from a model, writes the exact input and output token counts against the calling tenant, and answers 402 Payment Required once that tenant passes its monthly token cap. About 120 lines of Node. Thirty minutes, if Node is already on the machine.

Prerequisites:

  • Node 20 or newer — fetch is a global, so there is no HTTP client to install (Node globals).
  • An OpenAI API key from platform.openai.com/api-keys, on an account with billing enabled.
  • curl, or any tool that can POST JSON.

No Docker. No cloud account. No framework beyond Express.

What this is not: a UI, a signup flow, or a Stripe integration. It is the half that is painful to retrofit — identity, the usage record, and the enforcement point — with one AI feature bolted on top to prove the meter moves.

app.js, in full

Two tables. tenant holds the API key and the cap. usage_event is append-only, one row per model call, partitioned by YYYY-MM so a monthly total is one indexed SUM.

Quota check before the model call, usage write after the response lands. That order is not cosmetic. You can only bill tokens you actually spent, and a pre-flight token estimate is wrong on every request that isn't average.

// app.js — a metered AI SaaS endpoint. Node 20+, ESM.
import express from 'express';
import Database from 'better-sqlite3';

const PORT = Number(process.env.PORT ?? 8787);
const MODEL = process.env.MODEL ?? 'gpt-4o-mini';
const OPENAI_KEY = process.env.OPENAI_API_KEY;

if (!OPENAI_KEY) {
  console.error('OPENAI_API_KEY is not set. Get one at https://platform.openai.com/api-keys');
  process.exit(1);
}

const db = new Database('meter.db');
db.pragma('journal_mode = WAL');
db.exec(`
  CREATE TABLE IF NOT EXISTS tenant (
    id                TEXT PRIMARY KEY,
    api_key           TEXT UNIQUE NOT NULL,
    monthly_token_cap INTEGER NOT NULL
  );
  CREATE TABLE IF NOT EXISTS usage_event (
    id            INTEGER PRIMARY KEY AUTOINCREMENT,
    tenant_id     TEXT NOT NULL,
    period        TEXT NOT NULL,
    model         TEXT NOT NULL,
    input_tokens  INTEGER NOT NULL,
    output_tokens INTEGER NOT NULL,
    created_at    TEXT NOT NULL DEFAULT (datetime('now'))
  );
  CREATE INDEX IF NOT EXISTS usage_tenant_period ON usage_event (tenant_id, period);
  INSERT OR IGNORE INTO tenant (id, api_key, monthly_token_cap)
  VALUES ('acme', 'sk_test_acme_1', 100000);
`);

const q = {
  tenantByKey: db.prepare('SELECT * FROM tenant WHERE api_key = ?'),
  periodTotal: db.prepare(
    `SELECT COALESCE(SUM(input_tokens + output_tokens), 0) AS total
       FROM usage_event WHERE tenant_id = ? AND period = ?`),
  insertUsage: db.prepare(
    `INSERT INTO usage_event (tenant_id, period, model, input_tokens, output_tokens)
     VALUES (?, ?, ?, ?, ?)`)
};

const currentPeriod = () => new Date().toISOString().slice(0, 7); // e.g. 2025-11

const app = express();
app.use(express.json({ limit: '256kb' }));

// Every /v1 route is authenticated and attached to a tenant.
app.use('/v1', (req, res, next) => {
  const header = req.get('authorization') ?? '';
  const key = header.startsWith('Bearer ') ? header.slice(7).trim() : '';
  const tenant = key ? q.tenantByKey.get(key) : undefined;
  if (!tenant) return res.status(401).json({ error: 'unknown_api_key' });
  req.tenant = tenant;
  next();
});

app.get('/v1/usage', (req, res) => {
  const period = currentPeriod();
  const { total } = q.periodTotal.get(req.tenant.id, period);
  res.json({
    tenant: req.tenant.id,
    period,
    tokens_used: total,
    token_cap: req.tenant.monthly_token_cap,
    tokens_remaining: Math.max(0, req.tenant.monthly_token_cap - total)
  });
});

app.post('/v1/summarize', async (req, res) => {
  const text = typeof req.body?.text === 'string' ? req.body.text.trim() : '';
  if (text.length < 20) return res.status(400).json({ error: 'text_too_short' });

  const period = currentPeriod();
  const { total } = q.periodTotal.get(req.tenant.id, period);
  if (total >= req.tenant.monthly_token_cap) {
    return res.status(402).json({
      error: 'token_cap_reached',
      tokens_used: total,
      token_cap: req.tenant.monthly_token_cap
    });
  }

  const controller = new AbortController();
  const timer = setTimeout(() => controller.abort(), 20000);

  try {
    const upstream = await fetch('https://api.openai.com/v1/chat/completions', {
      method: 'POST',
      signal: controller.signal,
      headers: {
        authorization: `Bearer ${OPENAI_KEY}`,
        'content-type': 'application/json'
      },
      body: JSON.stringify({
        model: MODEL,
        max_completion_tokens: 200,
        messages: [
          { role: 'system',
            content: 'Summarize the support ticket in two sentences. Then a final line: SENTIMENT: positive|neutral|negative' },
          { role: 'user', content: text.slice(0, 8000) }
        ]
      })
    });

    const body = await upstream.json();
    if (!upstream.ok) {
      return res.status(502).json({ error: 'upstream_error', detail: body?.error?.message ?? null });
    }

    const inputTokens = body.usage?.prompt_tokens ?? 0;
    const outputTokens = body.usage?.completion_tokens ?? 0;
    q.insertUsage.run(req.tenant.id, period, MODEL, inputTokens, outputTokens);

    res.json({
      summary: body.choices?.[0]?.message?.content ?? '',
      model: MODEL,
      usage: { input_tokens: inputTokens, output_tokens: outputTokens },
      billed_to: req.tenant.id
    });
  } catch (err) {
    const aborted = err.name === 'AbortError';
    res.status(aborted ? 504 : 500)
       .json({ error: aborted ? 'upstream_timeout' : 'internal_error' });
  } finally {
    clearTimeout(timer);
  }
});

app.listen(PORT, () => console.log(`listening on http://localhost:${PORT}`));

prompt_tokens and completion_tokens come straight from the usage object the chat completions endpoint returns (API reference). That is the number the vendor bills you on, so that is the number you store. Character counts and request counts are proxies, and a proxy that disagrees with your invoice is a number you will end up reconciling by hand.

Build it

  1. Create the project: mkdir ticket-api && cd ticket-api && npm init -y && npm pkg set type=module. Check it took — cat package.json should now contain "type": "module".
  2. Install the two dependencies: npm install express better-sqlite3. npm finishes with a line like added 70 packages. If better-sqlite3 fails here, jump to "When it doesn't work" below.
  3. Save the file above as app.js in that directory.
  4. Export your key in the shell you will run the server from: export OPENAI_API_KEY=sk-... (PowerShell: $env:OPENAI_API_KEY="sk-..."). Do not put it in app.js.
  5. Start it: node app.js. You should see exactly listening on http://localhost:8787, and ls now shows meter.db, meter.db-shm and meter.db-wal — the two extra files are WAL mode doing its job.
  6. Confirm auth actually blocks. In a second terminal: curl -s localhost:8787/v1/usage returns {"error":"unknown_api_key"} with status 401. Add -i if you want to see the status line.
  7. Read the meter before spending anything: curl -s localhost:8787/v1/usage -H 'authorization: Bearer sk_test_acme_1'. You get {"tenant":"acme","period":"YYYY-MM","tokens_used":0,"token_cap":100000,"tokens_remaining":100000} with the current UTC month in period.
  8. Check input validation: curl -s localhost:8787/v1/summarize -H 'authorization: Bearer sk_test_acme_1' -H 'content-type: application/json' -d '{"text":"too short"}' returns {"error":"text_too_short"}, status 400, and no model call is made.
  9. Make the real call: curl -s localhost:8787/v1/summarize -H 'authorization: Bearer sk_test_acme_1' -H 'content-type: application/json' -d '{"text":"Ticket 4412 from dana@example.com: CSV exports have returned a 500 since Tuesday. She retried four times, sent two screenshots, and says she will cancel before renewal on the 30th."}'. You get back {"summary":"...","model":"gpt-4o-mini","usage":{"input_tokens":N,"output_tokens":M},"billed_to":"acme"}. The summary wording and the exact N and M will differ every run; what matters is that both numbers are present and non-zero.
  10. Prove the meter moved: run the /v1/usage call from step 7 again. tokens_used now equals input_tokens + output_tokens from the previous response, and tokens_remaining has dropped by the same amount. If it still reads 0, your usage write is not happening and nothing downstream of it can be trusted.
  11. Force the quota. Save setcap.js (below), run node setcap.js 10, then repeat the step 9 curl. You get {"error":"token_cap_reached","tokens_used":N,"token_cap":10} with status 402, and no request reaches OpenAI. Restore with node setcap.js 100000.
// setcap.js — change a tenant's monthly token cap.
import Database from 'better-sqlite3';

const cap = Number(process.argv[2]);
if (!Number.isInteger(cap) || cap < 0) {
  console.error('usage: node setcap.js <integer-cap>');
  process.exit(1);
}

const db = new Database('meter.db');
db.prepare('UPDATE tenant SET monthly_token_cap = ? WHERE id = ?').run(cap, 'acme');
console.log(db.prepare('SELECT id, monthly_token_cap FROM tenant').all());

node setcap.js 10 prints [ { id: 'acme', monthly_token_cap: 10 } ]. That is the entire plan-change mechanism, and it is enough to test enforcement without a billing provider in the loop.

Turning stored rows into money

At the listed gpt-4o-mini rates, the 100,000-token cap seeded into the tenant row is worth between roughly 1.5¢ (all input) and 6¢ (all output) of vendor cost. That is deliberately a test fixture: cheap enough to burn on purpose, shaped exactly like a real plan. Pricing the stored rows means pricing the two columns separately, because they differ by 4x:

SELECT tenant_id,
       model,
       SUM(input_tokens)  AS in_tok,
       SUM(output_tokens) AS out_tok,
       ROUND(SUM(input_tokens)  * 0.15 / 1000000.0
           + SUM(output_tokens) * 0.60 / 1000000.0, 4) AS usd_cost
  FROM usage_event
 WHERE period = '2025-11'
 GROUP BY tenant_id, model;

The two literals in that query are gpt-4o-mini rates only, which is precisely why usage_event carries a model column and why the grouping includes it. Once you serve a second model you replace the literals with a rate table keyed by (model, effective_date) and join on it. Cost is computed from the rows and never written into them: when a vendor changes a price, you re-run the query instead of migrating history. Append-only is what makes that safe — there is no destructive update anywhere in this schema, so every invoice you have ever sent can be recomputed from the same 100,000 rows that produced it.

Where the request goes, and what each failure returns

Request path: auth, quota gate, model call, usage write — and the status code each failure returnsPOST /v1/summarizeAPI key lookupQuota gateModel callWrite usage row200 + summary401 unknown key402 cap502 / 504upstream
StatusBodyCauseWho fixes it
401unknown_api_keyMissing or wrong Authorization: BearerCaller
400text_too_shortUnder 20 characters after trimCaller
402token_cap_reachedMonth's total is at or above the capCaller upgrades, or you raise the cap
502upstream_errorOpenAI returned non-2xx; detail carries their messageYou
504upstream_timeoutThe 20-second AbortController firedYou

402 is the reserved code for exactly this, per RFC 9110 §15.5.3. Do not use 429 here. A well-behaved client reads 429 as "back off and try again," so it will — every few seconds, for the rest of the month, against a budget that is not coming back.

Three things that break past one process

Two requests can both pass a gate that only had room for one. The quota check is a SELECT, and the matching usage write happens 1–20 seconds later. So N concurrent requests can all read total = cap - 1 and all proceed. The overshoot is bounded — max_completion_tokens is 200, so the worst case is N × (input + 200) tokens past the cap — but it is not zero. If that matters, reserve first: commit a row with the estimated input tokens and zero output before the await, then UPDATE that same row with the real counts when the response lands. better-sqlite3 transactions are synchronous and cannot wrap an await, which is why the reservation has to be its own committed write rather than one enclosing transaction.

SQLite has exactly one writer. WAL mode lets readers run concurrently with that writer, which is what db.pragma('journal_mode = WAL') buys and why you saw meter.db-wal appear in step 5 (SQLite WAL). One Node process on one box is fine. Two processes behind a load balancer sharing a network filesystem is not. At that point the meter moves to Postgres and the quota check becomes an atomic UPDATE tenant SET tokens_used = tokens_used + ? ... RETURNING against a counter row, so the read and the decision happen in one statement.

Streaming hides the usage object. Add stream: true and the token counts are not in the chunks by default; you have to request them with stream_options: {"include_usage": true}, which appends a final chunk carrying usage (API reference). Ship streaming without that flag and every metered row lands as 0 input, 0 output — the same silent failure described below, except now it arrives on the day you shipped a feature everyone liked.

When it doesn't work

better-sqlite3 refuses to load. The error is either Error: Could not locate the bindings file or a version mismatch: The module was compiled against a different Node.js version using NODE_MODULE_VERSION 115. This version of Node.js requires NODE_MODULE_VERSION 127. A native module built for one Node version, loaded by another. Usually right after nvm use. Fix: npm rebuild better-sqlite3, or rm -rf node_modules package-lock.json && npm install under the Node version you will actually run. See the better-sqlite3 repo for build requirements if the rebuild itself fails.

The key isn't reaching the process. If node app.js exits immediately with OPENAI_API_KEY is not set, you exported the variable in a different shell than the one running the server. If the server starts but /v1/summarize returns {"error":"upstream_error","detail":"Incorrect API key provided: sk-...","...":null}, the key is present but wrong or revoked. Verify it independently: curl https://api.openai.com/v1/models -H "authorization: Bearer $OPENAI_API_KEY" should return a JSON list, not an error. Rotate at platform.openai.com/api-keys.

The model name or a parameter is rejected. detail will read The model 'gpt-4o-mni' does not exist or you do not have access to it, or Unsupported parameter: 'max_tokens' is not supported with this model. Use 'max_completion_tokens' instead. The first is a typo, or an account without access to that model — override with MODEL=... node app.js. The second is why this code sends max_completion_tokens. Copy an older snippet that uses max_tokens and you get that error back instead of a summary (API reference).

One more, and it is the quiet one. If tokens_used stays at 0 after a successful summary, your usage_event insert is failing silently or the model returned no usage object. Invoices, quotas, margin per customer — everything on top of that row is then wrong, and wrong in the direction that costs you money.

The meter is a product decision, not plumbing

Your input cost scales with tokens, because that is the unit the vendor bills (OpenAI pricing). Per-seat pricing pretends otherwise. If usage varies 50x across your accounts, per-seat turns gross margin into a lottery you are not running.

Metering is routinely deferred until after launch, and the retrofit is expensive for a structural reason rather than a technical one: usage-based pricing needs a usage record that predates the pricing decision, and history you never wrote down cannot be backfilled. The usage_event table above is 7 columns and one index, and writing it on day one costs nothing; recovering three months of per-tenant token counts you never captured costs whatever the guess costs you. Build the meter first, then decide what you sell. That argument is developed at more length in start an AI company by building the meter first and how to make an AI startup, starting with the meter.

The obvious next move is turning /v1/summarize into something that picks its own next step — fetch the ticket, look up the customer, draft a reply, decide whether to send it. That changes the failure shape completely. A single call either returns or times out. A loop can spend without bound, and unbounded loops are the most-documented production failure in agent systems, which is why guidance on building them puts explicit stopping conditions and guardrails outside the model (Anthropic, Building effective agents). Where the loop decides is where your risk lives — that is the distinction in an AI agent decides its own next step.

So keep the two gates you just built in front of the loop, not inside the prompt. A step counter in your code and a token cap in your database cannot be talked out of firing. Granting an agent unsupervised spend authority on a live billing account leaves no enforcement point at all, and a model asked to police its own budget is not a control, because the thing being constrained is also the thing being asked.

Sources

  1. OpenAI API reference — Create chat completion — The usage object (prompt_tokens, completion_tokens) and the max_completion_tokens parameter used in app.js
  2. OpenAI API keys — Where the reader gets OPENAI_API_KEY
  3. OpenAI API pricing — Supports that API usage is billed per input and output token, so tokens are the billable unit
  4. better-sqlite3 — Synchronous SQLite driver used for the meter; prepared statements and install/rebuild behaviour
  5. SQLite Write-Ahead Logging — What journal_mode = WAL does and why -wal and -shm files appear
  6. Express 4 API — express.json — JSON body parsing and the limit option
  7. RFC 9110 §15.5.3 — 402 Payment Required — 402 is the reserved status code used for the quota response
  8. Node.js global fetch — fetch is available as a global in modern Node, so no HTTP client dependency is needed
  9. MDN — AbortController — The abort signal used to time out the model call
  10. Anthropic — Building effective agents — Agent loops need explicit stopping conditions and guardrails outside the model
  11. RFC 2606 — Reserved Top Level DNS Names — Why sample tickets use example.com addresses

Related