Semantic Models Explained: Why They Matter for Your Data & AI Strategy in 2026
What you'll learn
-
A semantic model is the connectivity and metadata layer that explains how data within a platform fits together. It is having a live moment: Databricks just released their version, Palantir has had one for years, and the debate over semantic models and ontology is all over LinkedIn.
-
The point of a semantic model is shared meaning. It ensures that a metric like profit means the same thing across the organisation and can be served consistently to whatever tool wants to consume it, whether that is an LLM or anything else.
-
Pointing an LLM straight at a SQL database and letting it riff produces answers that may not actually be accurate. The cautionary example is Databricks Genie, which happily answered questions over data sources but sometimes got the meaning wrong. A semantic model guides LLMs and other tooling so reports and outputs make sense.
-
Apache OSI is an open-source initiative for deriving semantic models that let you switch between providers or parts of your ecosystem while defining the data only once. There is support for it in dbt, Apache Polaris and elsewhere.
-
Semantic modelling should be a first-class consideration when spinning up data systems: how you model the data, and how you express that model so different tools can consume it. Saiku supports OSI models, with a couple of demonstrations available to experiment with.
By the end of this episode you should be able to explain what a semantic model is, why shared metric definitions matter before you deploy an LLM over your data, and where to start experimenting with OSI models.
In this episode
- What a semantic model actually is
- Why it matters for 2026: shared meaning across the organisation
- The Genie problem: LLMs riffing over raw SQL
- Apache OSI and portable, define-once models
- Getting started: dbt, Polaris and Saiku
A quick dive into semantic models, their growing importance in the data ecosystem, and how they're becoming essential for LLM deployment and organizational data consistency. Learn about recent developments from Databricks, Apache OSI, and how to get started with semantic modeling.
Show Notes
Key Topics Covered
What Are Semantic Models?
Definition and core concepts
Metadata and data connectivity within platforms
Ontology and data relationships
Why Semantic Models Matter in 2026
Ensuring consistent metric definitions across organizations
Guiding LLMs to provide accurate answers
Enabling data access control for different systems
Preventing AI hallucinations and inaccurate reporting
Industry Developments
Databricks: Recent semantic model release
Palantir: Long-established semantic model approach
Apache OSI (Open Semantic Initiative): Open source initiative for semantic model portability
Cross-platform data model interoperability
Real-World Challenges
Early LLM deployments over SQL databases
Databricks Genie accuracy issues
The importance of standardized metrics (e.g., 'profit' definitions)
Getting Started with Semantic Models
Tools and Platforms Mentioned:
Saiku Analysis Tool: demo.saiku.bi (includes OSI model examples)
dbt: Semantic model support
Apache Polaris: Semantic modeling capabilities
Resources
Saiku Demo: demo.saiku.bi
Apache OSI (Open Semantic Initiative)
Key Takeaways
Semantic models are essential for LLM accuracy and organizational data consistency
Open source initiatives like OSI are enabling cross-platform semantic model portability
Major players (Databricks, Palantir) are investing heavily in semantic modeling
Multiple open-source tools are available to start experimenting with semantic models
Next Steps
Explore semantic modeling in your data architecture
Test OSI models using available open-source tools
Consider how semantic models can improve your AI/LLM implementations
Chapters
0:02 - Introduction to Semantic Models
0:20 - Industry Developments: Databricks, Palantir & Apache OSI
1:00 - Why Semantic Models Matter in 2026
1:56 - The LLM Accuracy Problem
2:45 - Getting Started: Tools & Resources
Subscribe to our newsletter: https://newsletter.concepttocloud.com/
Want to apply AI to your engineering workflows? We build production ML pipelines, not demos.
Explore AI ServicesTranscript
Today we're gonna have a quick chat about semantic models, what they are, why you might want to consider them, and how you can get started. So [clears throat] if you've been following anything on LinkedIn, you'll have seen much chat and debate about semantic models, um, and also things like ontology. Databricks have just released their, uh, version of it. Palantir have had theirs for a long, long time. You know, and it's the connectivity of data within a platform, the, the metadata that explains how those things go together.
Now, [clears throat] the last couple of days has been, uh, in-- there's been in the news, uh, Apache Ossie, which is, uh, an open source initiative to be able to derive semantic models from different platforms to be able to, like, switch between different providers or different parts of your ecosystem, having only defined that data in one place. Why is this important? Well, because [clears throat] in, you know, twenty twenty-six, whilst you're trying your best to be able to deploy an LLM and have it answer the right questions, there's also the question of how you actually measure things with inside your organization, what those standards are, and how you give access to different pieces of data to different systems. And I don't mean just LLMs, I mean it could be anything. But you wanna be able to make sure that you measure it the same, so that profit means the same thing across your organization, you know?
And, [clears throat] um, and when you're doing that, you need a model that be-- is then capable of being able to provide that information to whatever tool then wants to consume it. Because at the end of the day, you don't really just wanna be able to s- deploy your LLM over your SQL database and have it just go and riff through some stuff and make it up. It's a little bit like when Databricks first released Genie, and it went and would happily answer questions over your data sources, but what did it actually mean, and was it actually accurate? And, well, sometimes it wasn't. Anyway, the, the, the idea with the, um, the semantic models is that you can define these things.
It helps guide LLMs. It helps guide other tooling to ensure that, you know, your reports and your output make sense. So why are we talking about this today? Well, like I said, it's been in the news because of Databricks, because of Ossie, and it's something that you should consider when you're spinning up, uh, data systems is, like, how do you model it, and then how do you expo- express that model in a way that different tooling can, uh, consume? If you wanna have a fiddle around with it, uh, we have support inside of our Saku analysis tool, which you can find a demo of on demo.
saku. bi. Um, and inside of there you can play around with an Ossie model, a couple of Ossie models, um, that demonstrate how that would work in an open source tooling environment. There is support for it, I believe, inside of DBT, Apache Polaris and elsewhere. Uh, so if you're interested in fiddling around with semantic models, that's as good a starting point as any.
I hope this has been useful. I hope that you can join me again next time for more information about how to leverage your data and AI. Thank you very much for joining me. I'll see you again soon. Why hire when you can partner?
Concept Cloud's leading engineers build your startup's prototype without the overhead. Launch faster. Conceptcloud. com.
Further reading
Why Your AI Project Is Actually a Data Project
The core thesis behind why a semantic layer, not the model, decides whether your LLM gives trustworthy answers.
AI Data Preparation & Agentic Workflows
The engineering seat for actually modelling data and expressing it so different tooling can consume it consistently.
Rebuilding Saiku With AI Agents
Background on the Saiku analysis tool referenced as a place to experiment with OSI semantic models.
More from The AI Briefing
AI Models Gone Rogue: OpenAI's ChatGPT Hacks Hugging Face & Security Implications
OpenAI's latest model attempted to hack Hugging Face instead of solving its assigned benchmark task. This episode explores the security implications of AI models exploiting vulnerabilities, the risks of open-weight models, and what businesses need to d...
SpaceX's Space Data Centers: The Multi-Trillion Dollar Gamble on Orbital AI
Tom explores Elon Musk and Sam Altman's recent Twitter exchange about SpaceX's ambitious plan to launch AI data centers into orbit. He breaks down the technical and economic challenges of space-based computing, from rocket reusability to the global chi...
AI Auditability: Why Explainability Matters in Regulated Industries
Exploring the critical challenge of AI explainability in regulated sectors. This episode dives into why organizations in finance, healthcare, and compliance-heavy industries must prioritize audit-proof AI workflows over pure optimization.