Skip to content
Eventual ConsistencyEpisode 6

Ethical Data Practices with Kevin Hartman — Part 1

Databricks VP Kevin Hartman explores open standards, responsible data science, and how community-driven innovation shapes ethical technology development.

27:59Full transcript below
KH

Kevin Hartman

VP of Analytics at Databricks

Many organizations understand the potential of their data but struggle to achieve actual business outcomes, especially when pursuing initiatives like generative AI. The core problem lies in data fragmented across silos, leading to slow access, unreliable insights, and inflated operational costs. This prevents leaders from making timely decisions and capitalizing on market opportunities.

Kevin Hartman, VP of Analytics at Databricks and a lecturer at UC Berkeley, shares his experience advising partners on how to build robust data platforms. He presents a structured approach for consolidating disparate data into a unified, governed lakehouse architecture. The discussion details how to transition from fragmented systems to a secure, performant environment capable of supporting advanced analytics and AI.

Hartman outlines specific strategies for data integration, ensuring governance with tools like Unity Catalog, and optimizing data for low-latency access. He also explores methods for secure data sharing through protocols like Delta Sharing and clean rooms, enabling cross-organizational collaboration while maintaining strict compliance. This conversation provides data leaders with a practical framework to transform their data estate into a genuine asset.

Key Takeaways

Generative AI relies on a structured, governed data foundation.

Many organizations rush to GenAI solutions, but true impact depends on first consolidating fragmented data into a unified platform. A lakehouse architecture, secured and managed by governance tools like Unity Catalog, provides the essential infrastructure. This foundational step ensures data quality, accessibility, and compliance, which are non-negotiable for reliable AI outcomes.

Demonstrate data platform value with targeted, iterative initiatives.

Building a comprehensive data platform requires significant investment, but delaying value delivery risks stakeholder buy-in. Instead, identify high-priority use cases that solve immediate business problems and deliver tangible results quickly. Iteratively implementing these solutions, gathering feedback, and showcasing ROI generates excitement and builds momentum for broader adoption across the organization.

AI-driven optimization is crucial for low-latency data access.

Addressing data latency goes beyond simply consolidating sources; it requires active performance management. Modern data platforms embed AI to automatically analyze usage patterns, optimize queries, and manage partitioning on the fly. This automation reduces manual tuning efforts and ensures that executives and analysts consistently receive fresh, high-performance data for critical decision-making.

Clean rooms facilitate secure, multi-party data collaboration without data movement.

Sharing sensitive data across departments or external partners often faces significant compliance and security hurdles. Clean rooms, powered by protocols like Delta Sharing, allow multiple parties to bring their data together in a controlled environment. This enables joint analysis and the creation of new insights from combined datasets, without any party’s raw data leaving its original secure location.

Related: CorrDyn helps organizations with data assessment and data engineering to build unified, performant data platforms. Our expertise also extends to ensuring data reliability and developing sound AI strategy that delivers measurable business value.

Full Transcript

Jason: Hi everyone, this is Jason, producer of Data BS. Welcome to part one of our interview with Kevin Hartman, head of partner solution architects in the Americas at Databricks and lecturer for data science at UC Berkeley. This conversation between James and Kevin was so good it ran for over an hour and a half and we just could not bear to leave so much on the cutting room floor. So we decided to split the episode into two parts. We really hope that you enjoy part one of this interview where they discuss collaboration, community, and best practices in data science. Here we go.

James Winegar: Welcome to Data BS, the show dedicated to tackling the big questions impacting the world of data and ML AI without any of the BS. My name is James Winegar. Each week I sit down with guests from across the data ecosystem to unpack how they’re shaping their businesses or the businesses of others through their real-world application of data engineering, ML AI, infrastructure, analytics, and more. No fluff, unfiltered, but slightly edited to remove noise. This is Data BS. Let’s get into it. All right, Kevin. So, tell me about you, who you are, and what do you do?

Kevin Hartman: Hi everybody. I’m Kevin Hartman. I work at Databricks as a partner solutions lead and advisor to our consulting and SI partners in the Americas. I lead a team of what we call partner solution architects that work with global and regional partners to develop GenAI solutions and analytics solutions, create champions of our platform, and solve some of the world’s toughest data problems.

James Winegar: What does it mean to lead a team in the partner solution architect space? What does that mean for your day to day?

Kevin Hartman: Day to day, I could dot in and get involved with numerous conversations with our partners. It typically involves enablement and some education, maybe review of solutions that run on our platform. To put this simply stated, my goal is to help our partners be better partners. My teammates — I manage a pretty small team, but mighty team, and each team member focuses on a different portfolio of regional and consulting partners to help them move forward. We’ve got a number of different initiatives, things that we seek to enable our partners with. I mentioned champions. One of the programs that we have is something called the Databricks partner champion. These are people who are well-versed in data and AI solution creation on our platform, well-versed in Databricks. They’re certified, they’ve done a number of solutions for customers already. They’re viewed as an extension of our own internal solution architects. It’s a really neat program, a really neat culture, and it’s wonderful to be part of that community.

James Winegar: When you talk about SI partners and things like that, you got your global regional boutique kind of firms that you work with. Your globals are your Accenture and Deloitte and all the other ones, your regionals, Lovelytics is probably a good example there.

Kevin Hartman: There actually is.

James Winegar: Yeah, you got your boutiques which is probably more like us, right?

Kevin Hartman: Uh, yeah, exactly. Excited to be talking to you about that too and hopefully one day we’ll be able to work on something together.

James Winegar: Yes. I don’t doubt it at all. The GenAI space is also interesting because you head up the GenAI portion of the partner solution architect organization, right?

Kevin Hartman: I do. I do focus a little bit on that, but I cross-cut. Our three main pillars that we are going to market with — there’s obviously generative AI. That’s really big. Something called Unity Catalog. We could talk about that a little later. And then just data warehousing. Those are the three big motions and uses of our platform that we are helping our partners with.

James Winegar: Yeah, I do focus a little bit on GenAI from my background, but there’s other pieces to that puzzle too.

Kevin Hartman: You got to build up to GenAI. You got to have your data together, supporting the data warehouse.

James Winegar: You have to. Those are the pieces you need to have put in place before you can really do some GenAI solutions that would be pretty impactful. Getting your data outside of silos and put into a unified platform for use.

Kevin Hartman: Let’s break those pillars out a little bit. If you had a fresh customer, my gut feeling would be that you’d go, let’s get your data together, let’s do this data warehouse thing. That’s going to be strictly focused on just getting it landed as an initial thing. Second thing is going to be how do we provide this out to the enterprise in a secure way that meets our compliance obligations and also enables better access and understanding of what is available using something like Unity Catalog. And then the third thing is, okay, now that you have these controls in place and you can support the various use cases — you have your data analysis, business intelligence, machine learning AI type use cases, and then GenAI I consider to be a little bit more full life cycle as compared to more traditional type of data workloads. As people evolve through that and leverage the platform, your team is focused on helping the solutions provider be as effective as they can within the ecosystem, right? It is an evolution of needs. Everyone wants to get to impact-driven insights as fast as they can. And it does take some steps to get your data organized and made accessible for some of these downstream solutions. The first thing you need to do is pull your data from where it may be now and get it into a unified platform like what we have. Organized and then governed by Unity Catalog, which is an open source technology. Those are some of the pre-steps so that you can do really cool analytics or insights or generative AI solutions. To make that easier, we have what we’re calling now a Data Intelligence Platform that combines aspects of GenAI out of the box to do insights on your data. But what you need to do is topology and organization, getting your data organized in a way that makes it understandable. Metadata on top of say your tables and your columns, so you have some descriptive information about what those things mean. Our platform then allows you to do things in native language to ask about something. It’s something we call BI, or we just call Genie. It allows you to ask questions to your data in natural language and report back in response and show visualizations about whatever you had queried. And it would show you the SQL behind it. So if you needed to tweak it, you can. It really lets people who don’t have that skill set interact with their data in a way that they couldn’t necessarily before. It’s really, really powerful. And then of course, you can extend that and make it even more powerful for your domain, for your data. We go into this philosophy of it’s your data. And it’s unique to your organization. It’s your organization’s investment, your IP. Now you need to look at how do I optimize that? How do I leverage that? How do I monetize that? How do I use that to get to things that I want to take to market or things that I want to do to be more efficient? There’s lots of ways you can start to use and leverage your data once it’s been collected and organized into a lakehouse architecture. Getting back to the original question you asked around the pillars. The use is one of the first things that you need to take that step into, but making sure that you have proper governance and organization. That’s the precursor to doing anything else. And that takes work sometimes. In a greenfield environment, it might be pretty easy, but that doesn’t exist. There’s usually some information that’s already in these silos or in legacy systems that need to be pulled into your lakehouse. I mentioned lakehouse a lot. That’s been an adopted standard now that Databricks invented several years ago, but now it’s been adopted as a strategy for many organizations that also have data platforms.

James Winegar: The first thing we talked about was you need to get your data landed into the lakehouse or whatever. How do you advise the approach for integration of these various disparate sources across silos?

Kevin Hartman: It’s integration or it’s migration. The first thing that you would want to do is take stock and run some assessment to understand what your needs are going to be and what you want to lift and shift versus what you just need to keep in place. There are different strategies for developing a solution once you have that pre-information. Planning is important. Once you have a plan, it’s then iterating on how you bring that data forward and also the strategies. There’s something called Lakehouse Federation. Going back to what I mentioned about Unity Catalog and why you want to use that as the underpinnings behind everything else, it really is a way to also govern all of your data and also AI assets too. You can govern models there. It also enables just-in-time pushdown controls — it’s a truly unified way of organizing and collecting your data. But you can also do federated connections to other data sources. Leaving that data in place, but still governing it and managing it as an asset under Unity Catalog. So you still have a unified place where you can manage permissions and things like that. Figuring out that data estate strategy is going to be your first thing. Once you have that down, you start iterating on that and going forward with maybe starting with a use case. Just taking what is something I want to deliver on once I have my data in a place where we can make it actionable. And then trying it out because you also need to consider people in this equation and your stakeholders and your proof of value. You can’t just go off and do something and then emerge six months later or maybe a year later, who knows. You said oh now I’m done with everything, now start to use it. That’s not how this works. There’s definitely things that may be driven by priorities, so you got to consider those first and take those high priority items that will deliver on most value and implement those first and get stakeholder feedback on that. Get buy-in and then once you get that, you’re generating excitement, you’re getting people to leverage and use. It starts to self-feed. If you are a director or manager owning this space, you have to show and report on value. Figuring out those use cases that are going to deliver on value would then enable you to keep going and eventually you develop a full ecosystem that has all of your data state laid out and working on your use cases that are really important to the business.

James Winegar: And then once you’re onboarded enough use cases, you have a whole process to run through. Now onboarding a new use case is a lot easier. Maybe you already have the data sources cataloged. So you already have the access controls within Unity Catalog. And then you can be like, okay, this use case, we can do without any onboarding steps because XYZ.

Kevin Hartman: It becomes self-service, right?

James Winegar: Yeah, and you don’t have to worry about getting data engineering or whoever to go in and figure out how to get on some Postgres database somewhere to federate the access across because it already happened at some point because it was, I don’t know, your bank and it’s your debit database or something. That wouldn’t be on Postgres, but you know what I’m going as far as.

Kevin Hartman: Yeah. Yeah.

James Winegar: One of the things that you said that triggered me, not triggered, but made me think about something is that you’ve federated this access to these other systems as far as where the data is, and/or you’re moving the data into the Databricks platform to support warehouse activities and things like that. Latency of that data movement, so how fresh do I need the data to be based on my use case, is a significant issue when I have these siloed data sets. How can Unity Catalog and these data governance tooling help support removing those barriers when you need low latency or always up-to-date data?

Kevin Hartman: Latency can be viewed in a number of different ways. There’s latency in just, I don’t even have that data. That’s the first level of latency that you need to address — making the data available. That’s the first thing that with Unity Catalog really enables you to get access in a very quick way and start discovering. Doing queries or doing natural language against your data. That’s one of the first tools that really helps solve the first level of latency problem. Then when you get to more things like, okay, well I solved that, but I need to consider that latency now as, okay, it’s connected, but a query is slow. I need this to be faster. That’s a whole other level of planning for and against your data estate. Now, if you’re built on a Lakehouse, there’s lots and lots of tools that are available to you to do really sophisticated query optimizations. Going back to what I mentioned around the platform itself and the use of generative AI, AI is embedded now into the platform, the platform being what we call now the Data Intelligence Platform. What it does is it inspects and analyzes the queries that you make, the things that you’re doing against your data and automatically optimizes things like partitioning and query generation based on that information. Re-optimizing on the fly based on your usage patterns. That’s another level of performance that will allow you to get to actual insights quicker. But then you’re now thinking about, well, what about data that I didn’t move into the Lakehouse and now I’ve got to work with a connected data source and it’s a different legacy warehouse system and I’m still federating it. That’s going to be a strategy where you need to consider, do I need to pre-optimize that? Are there certain queries that I just need to be bringing in and caching or making ready? You want to avoid if possible data replication. But if things are needed to be done in a very quick way or need insights in a very quick way, then you’re going to need to do some of that, at least maybe in an aggregate form. Or maybe you’re leveraging your connected data source to do the pre-aggregates for you. So you’re leveraging the power of that platform or that system to come up with whatever sums or aggregates that you need and then you’re just pulling that data forward. So it gets a little bit more complicated or potentially sophisticated depending on how you’re federating. But the good news there, the more you shift and move onto the Lakehouse platform itself, it’s becoming a solved problem as far as query performance is concerned.

James Winegar: I think there’s a story around change data capture CDC from your data sources. That’s going to be not zero milliseconds, but it’s going to be pretty fast to get the data replicated. You can do things like — it’s been a while since I did this in Databricks — but something like Autoloader etc. to get that data landed as quickly as possible. And then you can have Delta Live Tables see that change at the raw layer or the bronze layer and then push up through aggregates etc. in an automated fashion.

Kevin Hartman: All those things are sequenced and made ready for you in the lakehouse architecture you just mentioned. That’s called a medallion architecture you just mentioned. The whole philosophy behind that is, let’s get the data moved over quickly. The first thing you want to consider is just raw ingress, take it and land it first as your first stage. As part of that, you may do some minimal processing, but it’s largely kept fairly in its original state. This is important for data science work or any downstream work. Then you want to either enrich or refine. That’s usually done in what we call the silver layer, and then you are making the data ready for downstream consumption in what would be your gold layer, or your curated layer is another name for that. That’s going to feed into other systems that could go into a mart or could be a data science initiative as we mentioned before. The platform really allows you to go back to any state of that data as well, so you can see or look at and do analysis on what may have been in the silver layer instead. Putting all these pieces together, I think you mentioned this with Delta Live Tables. DLT is a framework, a declarative language that you can put into your definition of what you’re doing, your transformation. So when it comes to monitoring and scaling and running that query or that transformation in the environment, if you’re using DLT, those things become automatic, it becomes observable. All the information that is happening in real time gives you a view of how your pipelines are being managed. That’s another impactful thing that you can do on platform — leverage DLT to help with some of the scaling and performance things that you may have had to do on your own in the past or manually in the past.

James Winegar: My understanding is it’s very good for two windowed aggregates over some data, have that be updated in real time, do the partition pruning as it goes through etc. So that way you’re reducing that latency it takes to get these aggregate data sets or these final data sets that are supporting either a data science initiative or an analytics initiative, for an executive dashboard. And based on the use case, sometimes seconds matter, sometimes minutes are fine, sometimes hours are fine, and sometimes days are fine. Where’s the needle based on what you’re trying to do. It’s an optimization game. What are you going to pay for versus what you need.

Kevin Hartman: True. True. But the data intelligence engine now, it will auto-optimize based on your usage patterns too. If you’re running that query every five minutes or whatever it might be, it will start auto-optimizing for that usage pattern. If you’re using it only once a day, again, it really depends. Those are some of the things that become solved problems for you so you’re not having to go into the weeds and figure it out yourself.

James Winegar: Can you talk more about how data sharing in an organization can support either cross-team or cross-organization collaboration? Particularly in these large organizations where they have more compliance obligations that they have to manage and you don’t necessarily have a single production workspace that’s running.

Kevin Hartman: Oh gosh, that’s packed with a lot of different kinds of questions all in one because now security and compliance becomes an interesting dynamic to this too. Let’s start with just in general data sharing. I’ve probably mentioned — I think I mentioned this — that Databricks has a philosophy of following open standards, open formats, promoting and contributing to open source initiatives, or taking something that may have been developed internally and then making that into something that is open source. And Delta Sharing — there’s the Delta format first, that was an open sourced protocol from several years ago. But then came with that something called Delta Sharing. Building on the Delta format, also having a protocol for establishing the sharing of that data with other stakeholders, it could be inside or outside your organization. And making it so that that data is not moved. It’s not moved over. It’s accessed from where it is, but it facilitates the running of a query or some sort of compute against wherever that data lives already.

James Winegar: We’re not moving the data, so we can turn off access to the data so that we don’t have a potential breach if we want to remove access.

Kevin Hartman: Right, so we’re leaving it where it is. There’s Delta Sharing as the first thing. Underpinning that, you then have things like what we call Databricks Marketplace, which if you want to monetize your data, you can make that available on a marketplace for consumption. You can charge for that. You’re providing a service. Another use case, or maybe more in general terms, is taking something that may have been made available through Delta Sharing and putting that into something called a clean room. I can bring my data through Delta Sharing and then you could bring your data through Delta Sharing and put it into a clean room, so then you still have everything that is discoverable inside of that clean room through Unity Catalog exploration to start to do information gathering and queries across your data and my data and joining things and creating something that may be truly unique, but it stays within the confines of that clean room. Also, I can decide on what it is that I want to make available into that clean room too. I might want to say, okay, I need to scrub certain PII data — that could be something that you’re doing as a pre-step before sharing anything that could be sensitive. Or pre-agreeing on how you’re going to join it with some mechanism for creating that ID. Those are things that you can do with your data. Another interesting thing that you may want to do is not share your data directly but share the characteristics of your data. That’s where we were talking about synthetic data generation. There’s a whole other piece to that puzzle — what you could do as an organization is not provide the exact data that you might have but create synthetic alternatives that look like that and follow the same distribution. If you’re looking for not necessarily the data itself, but the characteristics instead. That’s something else that you can also make available and share either directly in the marketplace or in the clean room when you’re looking at doing joins across different organizational data.

James Winegar: I could go into many rabbit holes with the synthetic data thing. There’s a recent paper about how synthetic data sets are not actually producing model artifacts that are behaving as usable, because it’s generating distribution on a column level typically, but the distribution across the data set doesn’t make any sense in my opinion. Somebody could prove me wrong, but in a highly-dimensional data set it’s extremely hard.

Kevin Hartman: Of course. In high-dimensional data with many different features, that could be challenging. But for single column imputation type of thing — because a similar strategy can be employed for data imputation. I’m missing some fields, I’m missing something, so I need to generate something that is close enough.

James Winegar: The simpler the data is, the more realistic it is. The more complicated the data is, the harder it is to synthetically create something that matches appropriately. There was another thing you talked about with clean rooms. In my experience, clean rooms almost always have been for due diligence processes. Have you seen other use cases for clean rooms that are interesting that you could talk about?

Kevin Hartman: That’s great. Yes, due diligence and fact finding is definitely a really common place for this. But it could also be multi-organization. You’ve got data on consumer behavior, I’ve got data on advertisement campaigns, someone else has data on, I don’t know, buying patterns or something like that. There’s different things that can come together from different parties to build a different kind of data set.

James Winegar: Okay, so you got like Experian plus Nielsen plus your own data.

Kevin Hartman: Plus your own data, right? That’s the secret ingredient there. You’ve got source information based on the real promotional effects of something or campaign effects on something from a retailer. And yes, Nielsen is definitely one of those types of places that owns a lot of that data, Experian has a lot of that data. But you have your own actuals, like your sales on your product or inventory. The missing ingredient then is, okay, how do I more effectively run my own campaigns given what I see out there in the marketplace to make me more competitive?

Jason: That’s it for this episode of Data BS. If you enjoyed this episode, make sure you subscribe wherever you listen to your podcasts to not miss the next one. This episode was sponsored by CorrDyn, a data consultancy that helps organizations unlock the power of their data. If you have a data challenge, we can help. Visit corrdyn.com, c-o-r-r-d-y-n dot com to learn more. See you next time.

Frequently Asked
Questions

What's the best initial approach for bringing our disparate data sources into a unified platform?
Begin with a thorough assessment to understand your existing data estate and business needs. You'll then determine what data to lift and shift into the lakehouse versus what can remain in place through federated connections. Prioritize high-value use cases to demonstrate early impact and gain stakeholder buy-in.
How can Unity Catalog improve our data governance and access control across our data and AI assets?
Unity Catalog provides a single, unified place to govern all your data and AI assets, including models and external data sources via Lakehouse Federation. It enables centralized permission management and pushdown controls, ensuring consistent security and compliance across your entire data estate without moving data.
Our compliance team is concerned about sharing sensitive data with external partners. How can we enable collaboration securely?
Secure collaboration can be achieved using clean rooms, which allow multiple parties to combine and analyze data without any raw data leaving its original environment. Protocols like Delta Sharing facilitate this by providing controlled access to data where it resides. Additionally, consider using synthetic data for characteristic-based analysis where direct sensitive data sharing is too risky.

Ready to level up your data stack?

CorrDyn helps companies evaluate, build, and optimize their data platforms. The same team behind this show, working on your problems.

Book an intro call