OpenAI Select Partner
Union Square Consulting is an OpenAI Select Partner in the OpenAI Partner Network.
OpenAI Partner Network (opens in a new tab)Client-side through production
Three senior engineers working in the systems you already have: adoption and agent architecture, data and context, and the infrastructure beneath them.
Production AI is built inside your estate and carried through go-live. The result is designed for your team to run without a permanent consultancy retainer.
Shipped inside
Credentials
OpenAI has named us a Select Partner. Anthropic has published our Brainlabs work. We choose models on technical, operational and commercial fit.
Union Square Consulting is an OpenAI Select Partner in the OpenAI Partner Network.
OpenAI Partner Network (opens in a new tab)Third-party case
Anthropic published the Brainlabs case study, including the four-week rollout and 91% adoption in North America.
The published Anthropic customer story (opens in a new tab)3
Principals. One working unit.
2 wks
CFO concept to finance-grade production
40+ yrs
Combined engineering & delivery experience
~800
Processes automated, one estate
40
Subsidiaries, each on its own ERP, unifying onto one NetSuite
The practice
Three principals, one standard. The skills overlap on purpose, so the work is interchangeable. One of us leads on AI architecture. The other two come from corporate data and platform careers. You get that mix on day one, inside organisations that do not buy theatre. If the honest answer is "don't build it", that's the answer you get.
Coverage overlaps on purpose. Any of the three can pick up the others' work, so a stream never hangs off one person. You do not pay for a kickoff quarter of people learning each other's judgement.
The data and infrastructure careers came first. Chatbot work from 2022. GenAI from 2024, so the failure modes are already priced in: evaluation at Meta, agent estates at 1,000-person scale, and the platforms those systems sit on.
Meta, Spotify, UBS, Starling Bank, S&P Global, Sky, Société Générale, the European Investment Bank, Viasat. A broken pipeline has a real cost there. That is the bar we bring to an AI engagement.
What we do
A forward deployment unit
We embed in the estate, ship to production, and stay after go-live. Shared standards on day one. The people you meet are the people who build.
Traditional consultancies
Big-4 engagement
A partner pitches, juniors build, you get a roadmap. 6–18 months to production, if it gets there.
Staff augmentation
CVs added to your backlog. Output depends on your own architecture and management bandwidth.
The offer
It needs adoption, a data platform that can carry the work, and infrastructure that will keep running. We cover all three.
Agent estates with verifier gates, platform governance, and the enablement that makes adoption stick: from C-suite advisory through hands-on build. Your people end up authoring the automations, not renting them.
Recent: founded and now lead the architecture of a ~800 skill, 82-plugin platform; cut per-session token overhead 46%; shipped a production MCP server.
TM
Principal Architect
Memory and semantic layers that give production agents the business's own context, on data platforms engineered to carry them. Agents open a task already knowing the relevant conversations, not a blank prompt.
Recent: a 12× throughput re-architecture; GenAI evaluation pipelines at Meta; executive KPI rails at Spotify.
AA
Principal Engineer
Airflow orchestration, warehouse structure and CI/CD deepened to self-healing. The platform recovers without a war room.
Heritage: real-time fraud and machine learning at Starling Bank; streaming platforms at UBS, S&P Global and Sky; Basel III reporting across 13 Société Générale entities.
GM
Principal Data Engineer
Coverage map · no single point of failure
Scroll the table sideways for all three columns
| Capability | TM | AA | GM |
|---|---|---|---|
| Custom harness delivery | Core strength | Core strength | Core strength |
| LLM & agent systems | Core strength | Core strength | Core strength |
| RAG, memory & context engineering | Core strength | Core strength | Core strength |
| AI evaluation & verification | Core strength | Core strength | Core strength |
| Prompt engineering & skills authoring | Core strength | Working capability | Core strength |
| Claude Code & platform governance | Core strength | Working capability | Working capability |
| Data platforms & warehousing | Working capability | Core strength | Core strength |
| CI/CD, cloud & infrastructure | Working capability | Core strength | Core strength |
| React / front-end delivery | Core strength | Core strength | Core strength |
| C-suite & stakeholder management | Core strength | Core strength | Core strength |
The principals
Three principals, one standard. The skills overlap on purpose, so the work is interchangeable.
FDE · AI adoption, agent systems & Claude architecture
Claude Certified Architect (Professional) and Claude Certified Associate. Builds the infrastructure other people's AI runs on: from C-suite advisory down to the compiler, the MCP server, and the hooks that agents cannot talk their way past.
Verify on Credly
Stack: Python · Go · TypeScript/Node.js · multi-agent systems · RAG · MCP · Snowflake · BigQuery · Cloud Run · Claude/GPT/Gemini
FDE · data platform, context & memory architecture
The data platform and the context the agents stand on. Shipped inside Meta and Spotify. Nine years in the working unit.
Stack: Python (asyncio/FastAPI) · dbt · Airflow · Snowflake · BigQuery · Databricks/Spark · Beam/Scio · GCP/AWS/Azure · Terraform · React/Vue/Next.js · TypeScript
FDE · pipeline orchestration, CI/CD & data infrastructure
The orchestration and the infrastructure the estate recovers on. Shipped inside Starling, UBS, Sky and Société Générale. Nine years in the working unit.
Stack: Airflow · Python · Java · SQL/Postgres · Kafka/Flink/Spark · AWS/Azure/GCP · Snowflake-pattern warehousing · CI/CD · Docker · Teradata/Hadoop heritage
How we deploy
The same pattern used inside banks and platforms, and inside a 1,000-person agency. Compressed because the mistakes already happened somewhere else.
Sit inside your teams; map workflows, constraints, data and real AI maturity. You leave with a value-ranked use-case backlog you could put in front of the C-suite.
First working proof in production-shaped conditions, with evaluation criteria agreed before we build.
Verifier gates and evaluation frameworks go in before scale, along with access control, identity and governance, so none of it has to be retrofitted later.
Workshops, playbooks, reusable skills and local champions, so your own people compound the value instead of renting it.
Self-healing operations, with cost and latency headroom engineered before it's needed.
Next step
Bring one workflow you actually run. In 30 minutes we will say how we would take it to production, what we would not touch, and whether it is worth commissioning at all. If the honest answer is "don't build it", that is the answer you get.