ModernDataWork
An agentic information service of MyDataWork
How the data-worker community keeps tabs on what matters
Sign in / Subscribe

This week’s issue

Monday, Jul 20, 2026 · 14 items · new issue every Wednesday, 8am ET
Data engineering & the warehouse/lakehouse 3 items
Vendor newsData engineering & the warehouse/lakehouseDatabricks
Databricks Blog · Read source ↗

Glaspoort has implemented a continuous integration/continuous deployment (CI/CD) pattern for its Lakebase, allowing database branching similar to code repositories. This approach aims to streamline the deployment and management of their fiber infrastructure data in the Netherlands.

MyDataWork POV — Here's a reality check: Glaspoort's CI/CD pattern for Lakebase isn't just a technical feat—it's a wake-up call for data practitioners. While sales and marketing have their systems of record, our world is still cobbling together makeshift solutions. Glaspoort shows the potential when we treat data operations with the same rigor as code. But let's not kid ourselves—branching databases won't solve the core issue. Until we have a unified platform that ties data workflows directly to business outcomes, we're just rearranging deck chairs on the Titanic. The industry should take note: it's time for data work to get its own system of record.
Discussion happens on Reddit — no comments are hosted here.
Vendor newsData engineering & the warehouse/lakehousePostgresRedshiftKafkaDebeziumAirflow
r/dataengineering · Read source ↗

TouchBistro's insights team developed a data pipeline to transfer restaurant data from a sharded Postgres database into a Redshift data warehouse. The project utilized Kafka, Debezium, and Airflow to manage the data flow, addressing challenges in integrating Debezium with Redshift.

MyDataWork POV — Building a pipeline to handle millions of records daily is no small feat, especially when you're navigating the complexities of sharded databases and Redshift integrations. What stands out is the gap in practical guidance on using Debezium with Redshift. This isn't just a technical challenge; it's a reminder that the data community often lacks detailed, real-world playbooks for these integrations. The industry loves to talk about the power of CDC (Change Data Capture) but rarely provides the step-by-step breakdowns that teams like TouchBistro desperately need to turn theory into practice.
See how MyDataWork relates to this item
MyDataWork users can leverage their existing warehouse connections to automatically catalog outputs, like those from Debezium or Airflow, into Snowflake or Databricks, ensuring that lineage and freshness are tracked without needing additional tool-specific integrations. Explore MyDataWork ↗
Discussion happens on Reddit — no comments are hosted here.
Vendor newsData engineering & the warehouse/lakehouseAmazon
AWS Big Data Blog · Read source ↗

Alight Solutions successfully transitioned from managing their own Elasticsearch infrastructure to Amazon OpenSearch Service. This migration resulted in a 55% reduction in costs and saved about 2,000 hours annually in operational tasks, while also providing access to advanced observability features.

MyDataWork POV — Migrating to Amazon OpenSearch Service wasn't just a cost-saving maneuver for Alight Solutions; it was a strategic realignment of priorities. This story is a wake-up call for companies stuck in the inertia of legacy systems. While Alteryx customers ponder their next move post-acquisition, they should take note: the real challenge isn't just persistence but modernization. Operational overhead can suffocate innovation. The lesson here is clear—if your infrastructure is bogging you down, it's time to re-evaluate your tech stack and refocus on features that propel, not paralyze, your business.
Discussion happens on Reddit — no comments are hosted here.
BI & analytics tools 3 items
Vendor newsBI & analytics toolsSAP
SAP News (Data & Analytics) · Read source ↗

SAP Business AI's Q2 2026 release focuses on enabling businesses to accelerate decision-making and empower employees to concentrate on core activities. The release introduces features designed to enhance operational efficiency and strategic focus.

MyDataWork POV — SAP's latest AI release reads like a manifesto for speed and empowerment, but the devil is in the details. Faster decisions sound great, but without contextual understanding, you risk making the wrong ones faster. The reality is, AI can suggest paths, but it won't replace the nuanced judgment of a seasoned data professional. AI needs to be more than a speed booster; it must be a context enhancer. Businesses that recognize this will truly see their teams focus on what matters.
See how MyDataWork relates to this item
With MyDataWork's AI-powered Use Case Recommendations, users could leverage SAP's advancements to gain context-specific insights, ensuring that speed doesn't compromise decision quality. The platform's ability to identify automation candidates could further streamline workflow, aligning with SAP's empowerment goals. Explore MyDataWork ↗
Discussion happens on Reddit — no comments are hosted here.
Vendor newsBI & analytics toolsSAS
SAS Blogs · Read source ↗

SAS, marking its 50th anniversary, collaborates with partners to explore the future trajectory of AI and analytics. These partners bring unique insights from various industries, offering a glimpse into the evolving challenges and applications of analytics and AI.

MyDataWork POV — As SAS hits the big five-oh, the real question isn't just what's next for them, but what's next for the rest of us relying on their tools. Partner insights are invaluable, but here's the catch: they often miss the messy middle where AI actually meets real-world chaos. Predictive BI promises to cut the line between question and answer, but without understanding the gritty context of each business, it's like assembling IKEA furniture in the dark. We need partners who not only sell solutions but also illuminate the context layer—because that's where the magic, or the mess, really happens.
Discussion happens on Reddit — no comments are hosted here.
Vendor newsBI & analytics toolsSAS
SAS Blogs · Read source ↗

The DementAI team has leveraged SAS Viya's capabilities in AI, machine learning, governance, and decision-making to potentially identify Alzheimer's disease up to two years earlier than traditional methods. This approach ensures the necessary trust and oversight required in the healthcare sector.

MyDataWork POV — The promise of AI in early disease detection like Alzheimer's is undeniably attractive, but it's also a reminder of the stakes involved in healthcare AI. It's not just about the technology; it's about the trust in that technology. Identifying Alzheimer's earlier is a breakthrough, but one must ask: how are these predictions integrated into patient care pathways without risking over-reliance on machine output? The blend of governance with AI is not just a feature; it’s a necessity to ensure that these predictions don’t become yet another unchecked checkbox on a doctor’s digital dashboard.
See how MyDataWork relates to this item
MyDataWork's AI-powered features can assist healthcare teams by recommending reuse opportunities for data and identifying automation candidates, potentially streamlining processes and enhancing early detection efforts like those used by DementAI. Explore MyDataWork ↗
Discussion happens on Reddit — no comments are hosted here.
AI platforms — data science & ML 2 items
Vendor newsAI platforms — data science & MLDatabricks
Databricks Blog · Read source ↗

Databricks has announced advancements in scaling document classification to handle over 100,000 labels. This capability is designed to support large-scale production workloads, enabling businesses to manage and categorize vast amounts of freeform data efficiently.

MyDataWork POV — Document classification at this scale is a reminder that we’re still in the early innings of understanding how to manage freeform data. The real challenge isn't just handling the volume but making sure the labels actually mean something to the business. Without context, those 100,000 labels might as well be confetti. It’s a classic case of technology outpacing its practical application. We need to ask: does this scale-up genuinely solve a problem, or does it just showcase capability?
See how MyDataWork relates to this item
For MyDataWork users, connecting your warehouse means you can automatically catalog the outputs from tools like Databricks. This ensures that even large-scale classification efforts are tracked, providing a clear lineage and freshness status for each label. Explore MyDataWork ↗
Discussion happens on Reddit — no comments are hosted here.
Vendor newsAI platforms — data science & ML
SAS Blogs · Read source ↗

The log-KDE method addresses the issue of fitting kernel density estimates to strictly positive data, ensuring that the resulting density curve does not assign non-zero probability to negative values. This adjustment is crucial for accurately modeling quantities such as lengths and mass, which cannot be negative.

MyDataWork POV — Here's the thing about methods like log-KDE: they reveal the persistent disconnect between the polished statistical models and the gritty reality of data work. In theory, KDEs are elegant solutions for estimating data distributions. But in practice, analysts often face the messy task of adjusting these models to fit the constraints of real-world data. The log-KDE method is a reminder that while statistical purity is appealing, it's the pragmatic tweaks that make data tools genuinely useful. Yet, you'll rarely see these adjustments celebrated in the glossy pages of industry rankings. They live in the 'final_v3_USE_THIS_ONE.xlsx' world, not the polished diagrams.
Discussion happens on Reddit — no comments are hosted here.
Governance, catalog & semantic layer 3 items
Vendor newsGovernance, catalog & semantic layer
SAS Blogs · Read source ↗

The healthcare sector faces challenges such as rising costs, workforce shortages, and limited resources, prompting a push towards AI to improve efficiency and care outcomes. However, the success of AI in healthcare also hinges on trust and regulatory compliance, not just technological capabilities.

MyDataWork POV — AI in healthcare isn't just about clever algorithms or predictive models; it's about building trust in a highly regulated environment. The real challenge is not just ensuring AI systems are accurate, but making sure they are understood and trusted by both providers and patients. This means transparent data practices and clear governance structures are essential. AI can certainly improve outcomes, but without trust, it's just another piece of tech that collects dust. Healthcare AI must integrate into existing workflows, respecting the nuances of medical practice, not bulldoze through them.
Discussion happens on Reddit — no comments are hosted here.
Vendor newsGovernance, catalog & semantic layerDatabricks
Databricks Blog · Read source ↗

AI transparency involves revealing the data, model behavior, and decision-making processes of AI systems. It focuses on governance and explainability to ensure trust and accountability in AI applications.

MyDataWork POV — AI transparency is touted as the answer to trust issues in AI, but let's not kid ourselves. Governance and explainability are only half the battle. We need more than clear models; we need clear workflows. If AI systems are to be truly transparent, they must also map the human context around their use. Without this, transparency becomes another buzzword, leaving data workers to bridge the gap between AI outputs and practical business insights.
Discussion happens on Reddit — no comments are hosted here.
Vendor newsGovernance, catalog & semantic layer
SAS Blogs · Read source ↗

The debate in financial services around agentic AI centers on the balance between automation and human oversight. Industry leaders are exploring the limits of AI autonomy, focusing on integration with governance and human expertise.

MyDataWork POV — The financial sector is finally asking the right question about agentic AI: not just how autonomous can these systems be, but where should we draw the line? This isn't about throwing every task into the AI bucket and hoping for the best. It's about strategic integration — figuring out which decisions require a human touch and which can be safely automated. This isn't a problem technology alone can solve; it's a puzzle that requires understanding the interplay of systems, governance, and human expertise. The real challenge is in the nuance, not the novelty.
See how MyDataWork relates to this item
Agent Studio helps financial services teams precisely scope where AI should support human decision-making. By defining use cases and governance structures, it ensures AI deployments align with both automation needs and human oversight requirements. Explore MyDataWork ↗
Discussion happens on Reddit — no comments are hosted here.
Decision intelligence 3 items
Vendor newsDecision intelligence
Forrester Blogs · Read source ↗

The Product Information Management (PIM) field is rapidly transforming, driven by advancements in AI technology. Recent industry research and vendor briefings highlight this evolution, emphasizing the integration of AI to enhance data accuracy, accessibility, and utility in PIM systems.

MyDataWork POV — Everyone's buzzing about AI's impact on PIM, but let's not get lost in the hype. AI might streamline data management, but without a structured understanding of the workflows around that data, we're flying blind. The real evolution in PIM isn’t just about smarter systems; it’s about contextualizing how teams use and interact with product data daily. So, while AI is speeding up the data management process, it's the work context that will truly redefine PIM's role in enterprises. Data catalogs and AI might sketch the map, but only the work context can fill in the details.
Discussion happens on Reddit — no comments are hosted here.
Vendor newsDecision intelligenceSAS
SAS Blogs · Read source ↗

Rules-based decisioning engines automate marketing actions by executing predefined business rules, streamlining campaigns without manual segmentation. AI integration enhances these engines by enabling more dynamic and responsive decision-making.

MyDataWork POV — Rules-based decisioning engines have been the backbone of marketing automation, but they often feel like a relic of a bygone era. The introduction of AI into this mix is not just an upgrade; it’s a paradigm shift. AI doesn’t just follow rules—it learns, adapts, and predicts, turning static systems into dynamic partners. Yet, the real magic isn’t in the AI itself but in how marketers can now pivot their strategies in real-time, making campaigns more responsive and personalized. The challenge? Ensuring these smart engines don't run amok without human insight.
See how MyDataWork relates to this item
MyDataWork's AI-powered features can help marketers identify automation candidates within their existing use cases, streamlining the integration of AI into decisioning processes. This ensures that marketing strategies remain agile and responsive. Explore MyDataWork ↗
Discussion happens on Reddit — no comments are hosted here.
Vendor newsDecision intelligenceFivetran
Fivetran Blog · Read source ↗

Fivetran's blog discusses the benefits of an open data infrastructure, emphasizing its role in ensuring reliable and cost-effective support for future data use cases, particularly those involving AI.

MyDataWork POV — The allure of an open data infrastructure is undeniable, but let's not pretend it's a panacea. While Fivetran touts reliability and cost-efficiency, the real story lies in how this 'unified foundation' translates into day-to-day operations. Without a true system of record for data practitioners, even the most open infrastructure can leave analysts scrambling to connect the dots between disparate tools and outcomes. It's not just about the pipes; it's about the platform where data work actually lives and breathes.
Discussion happens on Reddit — no comments are hosted here.
© 2026 ModernDataWork — an agentic information service of MyDataWork. Editorial commentary is AI-generated from MyDataWork's perspective and clearly labeled as opinion. Sources are summarized and linked, never reproduced.