Skip to main content

4 posts tagged with "Data Integration"

Data integration patterns, ingestion strategies, and pipeline design

View All Tags

Clerk + Supabase RLS: Tenant Isolation

· 15 min read
Puneet Gupta
Founder, Supaflow

A user can sign in successfully and still see or change another tenant's data. Authentication proves who the user is; it does not tell Postgres which organization rows that user may access.

The risky shortcut is to trust an organization ID sent by the browser. A caller can change that value. Another common mistake is to use auth.uid(), which represents a Supabase Auth user UUID rather than Clerk's string user ID. Membership lookups inside RLS policies can also become recursive and slow.

This tutorial shows how to make the verified Clerk session token the root of the authorization decision:

  • Clerk authenticates the user and supplies the active organization context.
  • Supabase verifies the Clerk token and makes its claims available to Postgres.
  • Postgres derives the user and tenant from those claims.
  • Row-Level Security applies indexed, non-recursive policies to every query.

By the end, you will have a reusable schema, JWT helper functions, a controlled tenant-bootstrap function, explicit read/write policies, and tests for personal accounts, organizations, role boundaries, and cross-tenant attacks. The integration uses Clerk and Supabase's native third-party authentication—without a Clerk JWT template, a shared Supabase JWT secret, or auth.uid().

The complete runnable implementation is in the supaflow-labs/clerk-supabase-demo repository. The snippets below are intentionally small enough to study; use the repository migration and tests when building the complete example.

How to Connect a Local SQL Server with ngrok or bore

· 9 min read
Puneet Gupta
Founder, Supaflow

You want to try Supaflow against a SQL Server that runs on your own machine -- a developer install on your laptop or a server on your office network. There is no public IP, no port forwarding, and no VPN between that database and the cloud. A Supaflow-hosted Agent needs to reach the database over the network, so localhost in the datasource form will not work for this test.

A TCP tunnel is the fastest way to prove connectivity during a proof of concept (POC). This guide shows two temporary options: ngrok, the popular managed tunneling service, and bore, a minimal open-source alternative that needs no account. Both give you a public host and port that forward straight to your local SQL Server, and both plug into the Supaflow datasource form the same way.

For production pipelines, deploy a self-hosted Docker Agent on a stable host inside the same private network as SQL Server. The agent connects to SQL Server over the local network and polls Supaflow over outbound HTTPS, so the database port stays private. This removes the tunnel relay and changing public endpoint from the data path and gives long-running pipelines a predictable network path.

We Moved 26M Rows for $81. Fivetran Estimated It at $1.9K.

· 13 min read
Puneet Gupta
Founder, Supaflow

For a while, we used a simple line: why pay 5x more for Fivetran?

It was a good line. Easy to understand. Easy to remember.

Then we ran the numbers on a real high-volume Supaflow workspace and realized we were underselling it.

The difference was not 5x. It was more than 20x.

In one real Supaflow workspace, over a seven-day usage window from June 13 to June 19, 2026, we ran 453 jobs and moved 26,297,690 rows across 11,102 objects (individual tables and streams) and a broad source and destination test matrix. Supaflow used 26.98 credits to do it.

This was not a neat one-connector benchmark. The workspace had 30 source connections and 9 destination connections, including HubSpot, Salesforce, Oracle Transportation Management, PostgreSQL, SQL Server, Salesforce Marketing Cloud, SAP SuccessFactors, Google Analytics 4, Google Ads, Shopify, Stripe, SFTP file feeds, and Airtable. Those jobs loaded into Snowflake, S3 Data Lake, PostgreSQL, and SQL Server, and even pushed data back out through reverse ETL into Salesforce and Salesforce Marketing Cloud.

We also did not give the system ideal conditions. We ran the matrix concurrently to put real pressure on it: source APIs, destination writes, object-level orchestration, schema work, resets, full resyncs, and many jobs competing for runtime at once.

That is why I trust the result more. It is closer to a busy customer account than a polished demo where one fast source writes to one fast destination in isolation.

At the current Professional list price of $3 per credit, that workload is $80.95 in Supaflow usage.

Then we entered the same row count, 26,297,690, into Fivetran's 2026 Pricing Estimator as a Salesforce connection on the Standard plan for a 1-200 person company.

The estimate came back at $1,908.86 per month.

That is not 5x. It is more than 20x lower.

Oracle Transportation Management Integration: The Complete Guide

· 18 min read
Puneet Gupta
Founder, Supaflow

Oracle Transportation Management integration is one of the more deceptive data engineering problems in the logistics stack. OTM exposes a well-documented REST API and sync looks straightforward on paper -- until you run it against real data and discover that the metadata catalog returns 400 on half the tables you need, "empty" timestamps arrive as literal 0, and sync-mode responses silently truncate at 1 MB.

This guide walks through how OTM data integration actually works in production: the objects that matter, the five quirks that break naive pipelines, the right way to handle incremental sync with OTM's server-side clock, and how to move OTM data into Snowflake or any cloud warehouse reliably.