html.smolsites.dev
Smolhub architecture · current state

One client, two PracticeHub accounts, one safe analytics view.

Smolhub ingests PracticeHub records, keeps raw data in R2, tracks sync state in D1, computes monthly snapshots, and serves them through the analytics API and dashboards. The account-aware model prevents equal numeric IDs from being confused across accounts.

Updated 20 Aug 2026PuntVital case studyProduction deploy pending credentials

The system in one view

Each layer has a different job. Raw data is durable, D1 is control/configuration state, and processed JSON is the fast report read path.

1

PracticeHub

Separate API accounts for the same client. Each account has its own patient, appointment, type, payment, and referral-source IDs.

2

SyncWorkflow

Fetches paginated endpoints, resumes partial work, and writes account-owned raw Parquet to R2. D1 stores cursors and leases.

3

D1 control plane

Stores client settings, registered accounts, account-scoped referral/type configuration, sync status, and refresh state.

4

ProcessedDataWorkflow

Reads account-owned raw data, keeps source provenance, joins within account, resolves safe identities, and writes monthly NP JSON.

5

Analytics API + UI

Loads the month snapshot, rehydrates labels, and presents reports, scorecards, trends, and patient detail.

PuntVital’s account boundary

The numeric referral-source ID is only meaningful inside its PracticeHub account. It is not a global client ID.

What was wrong

The client API flattened both lists into one array. A lookup by referral ID 7 could therefore select the wrong label. The fix preserves ph_source=account:<id>, scopes writes to the current account, and filters the settings UI to that account.

Storage responsibilities

D1 · control and derived state

  • Registered PracticeHub accounts and current/historical roles
  • Referral and appointment-type configuration keyed by account
  • Sync cursors, partial status, leases, and report refresh state
  • Client-level metadata and access control

R2 · durable data and read cache

  • Raw PracticeHub endpoint Parquet
  • Account-owned increments for large endpoints
  • Processed monthly snapshots at {slug}/processed/{yyyy-mm}-new-patients.json
  • Other monthly analytics snapshots, including Meta data

KV · operational settings

  • Meta system-user tokens and operational settings
  • Paid-social and Google referral-source selections
  • Not the authoritative home for account-owned PracticeHub rows

Workflows · long-running work

  • SyncWorkflow: API → raw R2 plus D1 cursor state
  • ProcessedDataWorkflow: raw R2 → monthly report JSON
  • Partial syncs resume and merge; stale partials force a fresh rebuild

What happens when a report loads

1
Find the client and monthThe analytics API validates access and selects the requested calendar month.
2
Serve the cached snapshotR2 is the fast read path. The JSON contains cohort patients, counts, revenue, and referral identifiers.
3
Rehydrate current labelsAccount-scoped D1 mappings provide the display labels without treating equal IDs as globally unique.
4
Queue refresh when staleIf the snapshot version is old, the API queues a processed workflow. It does not block the report request.

Do we need to reprocess the cached snapshots?

Yes for PuntVital’s processed report snapshots; no raw re-import is expected.

The raw PracticeHub Parquet and D1 account data are already the source data. After the Worker deploy, reprocess PuntVital’s affected monthly NP snapshots so the cached JSON is regenerated using account-scoped mappings. Do not delete or overwrite raw data first.

Recommended sequence

  1. Authenticate Wrangler and deploy the merged Worker.
  2. Confirm both PuntVital accounts are present and their raw objects are complete.
  3. Queue a forced report-snapshot refresh for PuntVital.
  4. Verify historical and current months, especially months containing referral ID collisions.
  5. Check counts and labels before closing the production issue.

When a raw sync is needed

Only if an account’s raw Parquet is missing, incomplete, or its D1 sync status is still partial. For PuntVital, the historical window is configured through April 2026 and the Innate Solutions account continues from the current account boundary.

POST /api/data-ops/report-snapshot-refresh/puntvital { "months": ["2026-01", "2026-02"] }

Current production status

Code: merged. Build: passing. Deploy: blocked.

PR #291 is merged into main. The production build and test suite pass. The manual Cloudflare deployment failed because the configured API token is invalid or expired, so production still serves the previous unscoped client response. Re-authenticate Wrangler, deploy, then run the refresh sequence above.