ComparisonVERIFIED 2026-08-08
SouthBase vs Langfuse
The best open-source option in the category and, candidly, the architecture blueprint much of it followed — including us. MIT-licensed core, genuinely self-hostable, and priced sanely. Acquired by ClickHouse in January 2026.
Where Langfuse is better3 CONCESSIONS
Start with what they do well
They are genuinely open source and we are not
Langfuse's core is MIT and you can run the whole thing yourself, forever, with no vendor in the path. We ship an Apache-2.0 collector and a proprietary platform. If self-hosting the backend is a requirement, Langfuse wins and we are not a candidate.
Their pricing is excellent and we are matching it, not beating it
Core at $29/mo with 100k units and unlimited users is a strong offer. Our pricing shape is deliberately similar. We are not claiming to undercut them.
They are more mature than we are
Langfuse is a proven, widely deployed product with a large community. We are pre-launch.
Where we think they fall short
Our actual criticisms
Teams who like Langfuse's tracing but find that evaluation is a set of primitives they have to assemble, with no calibration or attribution layer on top.
- 01
Evaluation is primitives, not a trust layer
You can run LLM-as-judge scoring in Langfuse. What you cannot do is find out whether that judge agrees with your human raters, whether it has a position or verbosity bias, or how wide the confidence interval on a score actually is. That gap is the whole reason we exist.
- 02
No failure attribution
When a multi-agent trajectory fails, you get the trace and a score. Localizing the failure to a step, tool, prompt or memory write is still your job, done by reading.
- 03
Neutrality is now a fair question, not an accusation
ClickHouse acquired Langfuse in January 2026. Nothing has gone wrong. But 'is this tool's roadmap independent of a database vendor's interests' is a reasonable question to ask over a multi-year horizon, and it did not exist before.
Side by side8 ROWS
The matrix, wins and losses both
| Dimension | Langfuse | SouthBase |
|---|---|---|
| Self-host the backend | Yes — MIT core | No — cloud only |
| Judge calibration & bias diagnostics | Not offered | The core of our evaluation product (in development) |
| Failure attribution | Not offered | On the roadmap — not shipped |
| Claude token accounting | Input / output / cached | Four buckets, with cache-hit ratio and a thinking-token trap guard |
| Maturity | Proven at scale, large community | Pre-launch |
| Seat cost | Unlimited users from $29/mo (2 users free) | Unlimited seats on every tier |
| Wire format | OTel-native | OpenTelemetry / OpenInference. The same instrumentation points at any OTLP backend. |
| Agent runs in your infrastructure | Self-hosted deployment can; their cloud does not | Yes — Apache-2.0 collector, redaction before egress |
A marked cell is the stronger option on that row. We did not omit rows we lose.
Don't switch if…
When Langfuse is the right answer
If you need to self-host the backend, or you are happy assembling your own evaluation on top of solid tracing, stay on Langfuse. We are a better fit only if you want the trust layer above the traces and are willing to use a hosted platform to get it.
And if you do want to try ours
The integration is OpenTelemetry, so trying SouthBase does not mean removing Langfuse — you can point the same instrumentation at both and compare the output on your own traffic.
Other comparisons