<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="https://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>The Bit Bucket</title>
    <link>https://blog.greglow.com/</link>
    <description>Thoughts from Microsoft Data Platform MVP and RD – Dr Greg Low</description>
    <language>en</language>
    <generator>Hugo -- https://gohugo.io/</generator>

    
    <item>
      <title>Book Review: The Dead</title>
      <link>https://blog.greglow.com/2026/08/20/book-review-the-dead/</link>
      <guid>https://blog.greglow.com/2026/08/20/book-review-the-dead/</guid>
      <pubDate>Thu, 20 Aug 2026 00:00:00 AEST</pubDate>

      <description>I’ve spent some time lately reading some older and more famous books. One that seemed fascinating was a short story by James Joyce called The Dead .
Author James Joyce was an Irish novelist and poet. He was also well-known as a literary critic. Wikipedia says he is regarded as one of the most influential and important writers of the 20th century. I’ve spent time in Ireland over the years, and I’ve seen a lot of tributes to him both up in Dublin, and also around Kildare. He’s treated as somewhat of a national treasure.
</description>

      <content:encoded><![CDATA[
        
          <img src="https://blog.greglow.com/TheDead_BookCover.png" alt="cover image" /><br />
        
        <p>I’ve spent some time lately reading some older and more famous books. One that seemed fascinating was a short story by 





  <a href="https://en.wikipedia.org/wiki/James_Joyce">James Joyce</a>

 called 





  <a href="https://en.wikipedia.org/wiki/The_Dead_%28Joyce_short_story%29">The Dead</a>

.</p>
<h2 id="author">Author</h2>
<p><strong>James Joyce</strong> was an Irish novelist and poet. He was also well-known as a literary critic. Wikipedia says he is regarded as one of the most influential and important writers of the 20th century. I’ve spent time in Ireland over the years, and I’ve seen a lot of tributes to him both up in Dublin, and also around Kildare. He’s treated as somewhat of a national treasure.</p>
<h2 id="the-dead">The Dead</h2>
<p>I was fascinated to read this book as I had some limited time available, and it’s short. I decided to listen to it on Audible while driving around recently. It was only 1 hour and 25 minutes.</p>
<p>The story is about a man named Gabriel Conrol and his relationship with his family and friends. In particular, it spends the first half of the book dwelling on he and his wife going to a family Christmas party. (It’s held every year on Twelfth Night). During the party, he needs to make a speech and he is particularly worried about this. He’s a teacher and he’s worried that the people listening won’t get the academic references in his speech.</p>
<p>I loved the way that Gabriel talked about his book reviewing work, where he enjoyed receiving the books for review more than the paltry amount of money he got for reviewing them.</p>
<p>Eventually, the party winds down and he’s one of the last to leave. He and his wife Gretta are heading off to a hotel for what Gabriel thinks will be a passionately romantic night. Sadly, that’s not how it turns out. Gretta had listened to a song before they left the family, and it had reminded her of a long lost love. And she’d been thinking about him ever since. Worse, she blamed herself for him dying because he came out in the lousy weather to see her.</p>
<p>After Gretta falls asleep crying, Gabriel is left thinking about the fact that there was this whole aspect of Gretta’s life that he knew nothing about, and thinking about how many dead people are living in the minds of people who still think of them. He realizes how he will also just be a memory one day, if remembered at all.</p>
<h2 id="summary">Summary</h2>
<p>I loved this book, mostly because I’ve also spent a lot of time over the years, pondering life and death, and coming to many of the same conclusions as Joyce through Gabriel.</p>
<p>I also enjoyed listening to the beautiful Irish accent and many (now-dated) Irish words from the time.</p>
<p>It was a great way to spend an hour and a half. I need to read more of Joyce’s work.</p>
<p>8 out of 10</p>

      ]]></content:encoded>
    </item>
    
    <item>
      <title>Writing SQL Queries for DuckDB Course Released</title>
      <link>https://blog.greglow.com/2026/08/20/writing-sql-queries-for-duckdb-course-released/</link>
      <guid>https://blog.greglow.com/2026/08/20/writing-sql-queries-for-duckdb-course-released/</guid>
      <pubDate>Thu, 20 Aug 2026 00:00:00 AEST</pubDate>

      <description>Even more SQL love !
Another popular SQL dialect is for DuckDB. And we’ve just completed our first course using it.
Creating reports, analytics, or applications? And need to get data out of DuckDB? Or are you building applications using DuckDB.
Learn to write SQL queries like a pro !
We now have very popular SQL courses, for all the most important SQL dialects. There’s our flagship course for T-SQL, then already courses for PostgreSQL, Snowflake, Oracle, DB2, MySQL, Azure HorizonDB, and Google BigQuery. We’ve just added our new course Writing SQL Queries for DuckDB and you can enrol in it now. It’s just $95 USD.
</description>

      <content:encoded><![CDATA[
        
          <img src="https://blog.greglow.com/DDS_Course_Advert_Banner.png" alt="cover image" /><br />
        
        <p>Even more SQL love !</p>
<p>Another popular SQL dialect is for DuckDB. And we’ve just completed our first course using it.</p>
<p>Creating reports, analytics, or applications? And need to get data out of DuckDB? Or are you building applications using DuckDB.</p>
<p>Learn to write SQL queries like a pro !</p>
<p>We now have very popular SQL courses, for all the most important SQL dialects. There’s our flagship course for T-SQL, then already courses for PostgreSQL, Snowflake, Oracle, DB2, MySQL, Azure HorizonDB, and Google BigQuery. We’ve just added our new course <strong>Writing SQL Queries for DuckDB</strong> and you can enrol in it now. It’s just $95 USD.</p>
<p>Check it out and enrol now here: 





  <a href="https://sqldownunder.com/courses/dds">Writing SQL Queries for DuckDB</a>

</p>
<h2 id="course-summary">Course Summary</h2>
<p><strong>Do you need to learn how to write SQL queries for DuckDB?</strong></p>
<ul>
<li>You know that the information that you need is stored in a DuckDB database</li>
<li>You need to find some data to extract it</li>
<li>You need to create reports with reporting tools</li>
<li>You are building applications using DuckDB</li>
<li>You are using analytic tools like Looker, Power BI, Tableau, QlikView, Excel, or Access and need to get data from DuckDB</li>
<li>You want to learn to write queries properly, using commercial coding standards</li>
<li>You are new to writing queries, or you are self-taught and want to make sure you are doing it correctly</li>
</ul>
<p>If so, this course is for you! And as well as detailed instruction, the course also offers optional practical exercises and quizzes to reinforce your learning.</p>
<p>We encourage you to complete the practical exercises. We have tried to make this as easy as possible. We’ve made setting up for the exercises in this course, both easy and free.</p>
<p>Check it out and enrol now here: 





  <a href="https://sqldownunder.com/courses/dds">Writing SQL Queries for DuckDB</a>

</p>

      ]]></content:encoded>
    </item>
    
    <item>
      <title>Fabric RTI 101: Governance and Compliance</title>
      <link>https://blog.greglow.com/2026/08/19/fabric-rti-101-governance-and-compliance/</link>
      <guid>https://blog.greglow.com/2026/08/19/fabric-rti-101-governance-and-compliance/</guid>
      <pubDate>Wed, 19 Aug 2026 00:00:00 AEST</pubDate>

      <description>Governance and compliance are critical considerations in real-time systems — they apply every bit as much as they do in traditional batch workloads. The difference is that in streaming environments, everything happens continuously, so your controls and monitoring must operate in real time as well.
The first step is to track data lineage — understanding where each event originated, how it was transformed, and what downstream systems or actions it affected. Lineage visibility helps you answer questions like Where did this number come from? or Which events led to this automated decision? It also supports root cause analysis when problems occur.
</description>

      <content:encoded><![CDATA[
        
          <img src="https://blog.greglow.com/FabricRTI101.png" alt="cover image" /><br />
        
        <p>Governance and compliance are critical considerations in real-time systems — they apply every bit as much as they do in traditional batch workloads. The difference is that in streaming environments, everything happens continuously, so your controls and monitoring must operate in real time as well.</p>
<p>The first step is to track data lineage — understanding where each event originated, how it was transformed, and what downstream systems or actions it affected. Lineage visibility helps you answer questions like <em>Where did this number come from?</em> or <em>Which events led to this automated decision?</em> It also supports root cause analysis when problems occur.</p>
<p>Next, ensure that all actions triggered from events are properly audited. When a real-time pipeline automatically takes an action — such as sending a notification, updating a database, or executing a workflow — those activities should be logged.</p>
<p><img src="https://greglow.blob.core.windows.net/blog/images/FabricRTI101_10_09_01.png" alt="Governance and Compliance"></p>
<p>Audit trails are vital for both operational transparency and regulatory oversight.
You should also apply compliance rules that match your organization’s or region’s regulatory obligations. Examples include GDPR for data privacy, HIPAA for healthcare data, and SOX for financial accountability. These frameworks may require you to manage where data is stored, how long it’s kept, and who can access it.</p>
<p>Related to that, make sure retention and deletion policies are enforced automatically. Streaming data can accumulate quickly, so your pipelines should include rules to archive or delete data according to retention limits. This not only supports compliance but also keeps costs and storage use under control.</p>
<p>Finally, implement tools that monitor for sensitive data in streams. That includes personally identifiable information (PII), financial details, or health data that may need masking or restricted handling. By flagging sensitive content early, you reduce the risk of data leaks or noncompliance.</p>
<p>Governance and compliance in real-time systems mean traceability, accountability, and protection — ensuring that every piece of data and every automated action can be understood, justified, and managed within regulatory boundaries.</p>
<h2 id="learn-more-about-fabric-rti">Learn more about Fabric RTI</h2>
<p>If you really want to learn about RTI right now, we have an online on-demand course that you can enrol in, right now. You’ll find it at 





  <a href="https://sqldownunder.com/courses/rti">Mastering Microsoft Fabric Real-Time Intelligence</a>

</p>

      ]]></content:encoded>
    </item>
    
    <item>
      <title>SQL: Using Database Snapshots to Provide Large Unit Testing Data with Quick Restores</title>
      <link>https://blog.greglow.com/2026/08/18/sql-using-database-snapshots-to-provide-large-unit-testing-data-with-quick-restores/</link>
      <guid>https://blog.greglow.com/2026/08/18/sql-using-database-snapshots-to-provide-large-unit-testing-data-with-quick-restores/</guid>
      <pubDate>Tue, 18 Aug 2026 00:00:00 AEST</pubDate>

      <description>SQL Server databases have a reputation for being hard to test, or at least hard to test appropriately.
For good testing, and particularly for unit tests, you really want the following:
Database in a known state before each test Database containing large amounts of (preferably masked) data (production-sized) Quick restore after each test before the next test For most databases, this is hard to achieve. The restore after each test means that a normal database restore can’t be used. What I often see instead, is people using transactions to try to achieve this i.e. the process becomes:
</description>

      <content:encoded><![CDATA[
        
          <img src="https://blog.greglow.com/DuplicateClocks_AC.png" alt="cover image" /><br />
        
        <p>SQL Server databases have a reputation for being hard to test, or at least hard to test appropriately.</p>
<p>For good testing, and particularly for unit tests, you really want the following:</p>
<ul>
<li>Database in a known state before each test</li>
<li>Database containing large amounts of (preferably masked) data (production-sized)</li>
<li>Quick restore after each test before the next test</li>
</ul>
<p>For most databases, this is hard to achieve. The restore after each test means that a normal database restore can’t be used. What I often see instead, is people using transactions to try to achieve this i.e. the process becomes:</p>
<ul>
<li>Start a transaction</li>
<li>Run the test</li>
<li>Check the results</li>
<li>Roll back the transaction</li>
</ul>
<p>In some situations, that works well but the minute that you start trying to test transactional code, things fall apart quickly. SQL Server doesn’t support truly nested transactions. When you execute a ROLLBACK, it doesn’t matter how deep this occurs, the outer transaction is being rolled back too.</p>
<p>One option that I’m often surprised that people don’t consider is database snapshots.</p>
<p><img src="https://greglow.blob.core.windows.net/blog/images/Snapshot.png" alt=""></p>
<p>Your testing mechanism becomes:</p>
<ul>
<li>Create a snapshot of the database to be tested</li>
<li>Test against the original database</li>
<li>Check the results</li>
<li>Revert the database to the snapshot</li>
</ul>
<p>Both creating a snapshot database, and restoring from the snapshot are very quick operations. The creation is always quick and the revert time depends upon how many pages were changed during the test. That’s often not many.</p>
<p>Now this does assume that you have your own database for doing this work, and that no-one else is connected to it. And no, there is no option to do this with Azure SQL Database (wish there was).</p>
<p>Creating and reverting snapshots is described 





  <a href="https://learn.microsoft.com/en-us/sql/relational-databases/databases/create-a-database-snapshot-transact-sql?view=sql-server-ver17">here</a>

.</p>

      ]]></content:encoded>
    </item>
    
    <item>
      <title>Fabric RTI 101: Security and Governance for Pipelines</title>
      <link>https://blog.greglow.com/2026/08/17/fabric-rti-101-security-and-governance-for-pipelines/</link>
      <guid>https://blog.greglow.com/2026/08/17/fabric-rti-101-security-and-governance-for-pipelines/</guid>
      <pubDate>Mon, 17 Aug 2026 00:00:00 AEST</pubDate>

      <description>Security and governance are just as important in real-time streaming systems as they are in traditional batch pipelines — and in some cases, even more so. Because streaming data is continuous, a single misconfiguration can expose sensitive data for long periods before it’s noticed.
The first principle is to secure event sources and destinations.
Every data connection — whether it’s an IoT device, API, or event broker — should use authentication to verify its identity and encryption in transit to protect data as it moves. In Microsoft Fabric, this typically means using TLS for secure transmission and managed identities for authentication, avoiding hard-coded credentials wherever they are available.
</description>

      <content:encoded><![CDATA[
        
          <img src="https://blog.greglow.com/FabricRTI101.png" alt="cover image" /><br />
        
        <p>Security and governance are just as important in real-time streaming systems as they are in traditional batch pipelines — and in some cases, even more so. Because streaming data is continuous, a single misconfiguration can expose sensitive data for long periods before it’s noticed.</p>
<p>The first principle is to secure event sources and destinations.</p>
<p><img src="https://greglow.blob.core.windows.net/blog/images/FabricRTI101_10_08_01.png" alt="Security and Governance for Pipelines"></p>
<p>Every data connection — whether it’s an IoT device, API, or event broker — should use authentication to verify its identity and encryption in transit to protect data as it moves. In Microsoft Fabric, this typically means using TLS for secure transmission and managed identities for authentication, avoiding hard-coded credentials wherever they are available.</p>
<p>Next, implement role-based access control (RBAC) within Fabric. RBAC ensures that users and services have only the permissions they need — for example, allowing operators to monitor pipelines without granting them the ability to modify configurations. This principle of least privilege helps limit the potential impact of errors or breaches.</p>
<p>Governance is another key part of pipeline management. Maintain visibility into schema evolution — knowing when and how data structures change — and track data lineage so you can trace where information originated and how it was transformed. This is critical for both regulatory compliance and operational troubleshooting.</p>
<p>Because streaming data often includes personally identifiable information (PII) or other sensitive details, it’s important to monitor for data exfiltration — whether accidental or intentional. This includes detecting unauthorized destinations or unusually high outbound data volume.</p>
<p>Secure pipelines combine technical controls (like encryption and RBAC) with governance practices (like lineage tracking and monitoring). Together, they ensure that your real-time system remains both trustworthy and compliant as it scales.</p>
<h2 id="learn-more-about-fabric-rti">Learn more about Fabric RTI</h2>
<p>If you really want to learn about RTI right now, we have an online on-demand course that you can enrol in, right now. You’ll find it at 





  <a href="https://sqldownunder.com/courses/rti">Mastering Microsoft Fabric Real-Time Intelligence</a>

</p>

      ]]></content:encoded>
    </item>
    
    <item>
      <title>Opinion: Don't just hire clones of yourself</title>
      <link>https://blog.greglow.com/2026/08/16/opinion-dont-just-hire-clones-of-yourself/</link>
      <guid>https://blog.greglow.com/2026/08/16/opinion-dont-just-hire-clones-of-yourself/</guid>
      <pubDate>Sun, 16 Aug 2026 00:00:00 AEST</pubDate>

      <description>Many years back, I was invited to chair a course accreditation panel for a local TAFE (Technical and Further Education) course. They had started to offer a computing-related 3-year diploma, and the hope was that it wasn’t too far below the 3-year degrees offered at local universities. One part of that accreditation process involved me discussing the course with the staff members who were teaching it.
After talking to almost all the staff, what struck me was how similar they all were. In the requirements for the course, there was a standard that each staff member needed to meet, but there was also a requirement for the group of staff to be diverse enough to have broad knowledge of the industry. There was no individual staff member that you could identify as not being at the appropriate standard, but almost all of them had exactly the same background, career progression, etc.
</description>

      <content:encoded><![CDATA[
        
          <img src="https://blog.greglow.com/Clones_FeaturedImage.png" alt="cover image" /><br />
        
        <p>Many years back, I was invited to chair a course accreditation panel for a local TAFE (Technical and Further Education) course. They had started to offer a computing-related 3-year diploma, and the hope was that it wasn’t too far below the 3-year degrees offered at local universities. One part of that accreditation process involved me discussing the course with the staff members who were teaching it.</p>
<p>After talking to almost all the staff, what struck me was how similar they all were. In the requirements for the course, there was a standard that each staff member needed to meet, but there was also a requirement for the group of staff to be diverse enough to have broad knowledge of the industry. <strong>There was no individual staff member that you could identify as not being at the appropriate standard, but almost all of them had exactly the same background, career progression, etc.</strong></p>
<p>The manager had basically hired clones of himself. It’s an easy mistake to make. If you feel you are the right person for a particular type of job, then hiring more people like yourself must help right?</p>
<h2 id="medical-research">Medical Research</h2>
<p>A similar problem happens in areas like medical research. Taking a whole bunch of people with the same background and experience isn’t going to let you cut through tricky problems that need someone to think outside the box.</p>
<p>Adding someone like a civil engineer into the mix might seem odd at first, but can have surprising outcomes. At the very least, they might ask a question that leads someone else in the team to think differently.</p>
<h2 id="application-developers">Application Developers</h2>
<p>I’m remembering this story because I see the same issue in application development groups.</p>
<p>I’ve done some work at a company that has over 400 developers. Data is almost all that they do, yet for most of the time the company has existed; they’ve had no-one focused on data. Everyone involved in development has a similar development background. They had many intractable data-related problems yet more and more of the same type of people was never going to solve those.</p>
<p><strong>Hiring a team of people who think and work like you do might seem like a good idea but it’s not. You need a mixture of people if you want to be really effective</strong>. (And that also means having a degree of gender and cultural diversity too).</p>

      ]]></content:encoded>
    </item>
    
    <item>
      <title>Fabric RTI 101: Replay and Reprocessing</title>
      <link>https://blog.greglow.com/2026/08/15/fabric-rti-101-replay-and-reprocessing/</link>
      <guid>https://blog.greglow.com/2026/08/15/fabric-rti-101-replay-and-reprocessing/</guid>
      <pubDate>Sat, 15 Aug 2026 00:00:00 AEST</pubDate>

      <description>In a real-time data system, it’s not enough to process events once and move on. There are many cases where you need to replay or reprocess event data that has already passed through the system.
Replay and reprocessing allow you to go back in time — to re-run data through your pipelines as if it were arriving again in real time.
This capability is especially useful for debugging issues, conducting compliance audits, or retraining machine learning models with historical data.
</description>

      <content:encoded><![CDATA[
        
          <img src="https://blog.greglow.com/FabricRTI101.png" alt="cover image" /><br />
        
        <p>In a real-time data system, it’s not enough to process events once and move on. There are many cases where you need to replay or reprocess event data that has already passed through the system.</p>
<p>Replay and reprocessing allow you to go back in time — to re-run data through your pipelines as if it were arriving again in real time.</p>
<p><img src="https://greglow.blob.core.windows.net/blog/images/FabricRTI101_10_07_01.png" alt="Replay and Reprocessing"></p>
<p>This capability is especially useful for debugging issues, conducting compliance audits, or retraining machine learning models with historical data.</p>
<p>For example, imagine that you’ve discovered a logic error in one of your transformations. Once you’ve fixed it, you can replay the stored event stream through the corrected logic to produce accurate results. Similarly, if a regulatory audit requires proof of how a system behaved at a specific time, replaying the original event data provides verifiable evidence.</p>
<p>In Microsoft Fabric, the recommended approach is to persist event streams into OneLake. By saving raw events in a durable storage layer, you can reprocess them later without relying on the original event source. This persistence also provides the foundation for combining real-time data with historical context.</p>
<p>Delta Lake plays a key role here. Its <strong>time travel</strong> feature allows you to query previous versions of tables, making it possible to analyze or reprocess data as it existed at any given point in time. That’s particularly valuable for iterative development and for maintaining reproducibility in analytics.</p>
<p>Replay and reprocessing effectively bridge the gap between streaming and batch analytics. They give you the flexibility to correct, retrain, or reanalyze data using the same infrastructure — ensuring that real-time systems remain not just fast, but also auditable, reliable, and adaptable over time.</p>
<h2 id="learn-more-about-fabric-rti">Learn more about Fabric RTI</h2>
<p>If you really want to learn about RTI right now, we have an online on-demand course that you can enrol in, right now. You’ll find it at 





  <a href="https://sqldownunder.com/courses/rti">Mastering Microsoft Fabric Real-Time Intelligence</a>

</p>

      ]]></content:encoded>
    </item>
    
    <item>
      <title>Book Review: Causal Inference with Bayesian Networks</title>
      <link>https://blog.greglow.com/2026/08/14/book-review-causal-inference-with-bayesian-networks/</link>
      <guid>https://blog.greglow.com/2026/08/14/book-review-causal-inference-with-bayesian-networks/</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 AEST</pubDate>

      <description>I recently received a review copy of Causal Inference with Bayesian Networks by Yousri El Fattah and Reza Bagheri from my friends at PackT.
Authors Yousri El Fattah is the CEO of Causal Computing and an expert in machine intelligence, causal modelling, control systems engineering, and data science.
Reza Bagheri is a working data scientist at Ipsos. He has written extensively on data science and machine learning, and has spoken at substantial conferences.
</description>

      <content:encoded><![CDATA[
        
          <img src="https://blog.greglow.com/CausalInferenceWithBayesianNetworks_BookCover.png" alt="cover image" /><br />
        
        <p>I recently received a review copy of 





  <a href="https://www.packtpub.com/en-au/product/causal-inference-with-bayesian-networks-9781835089217">Causal Inference with Bayesian Networks</a>

 by Yousri El Fattah and Reza Bagheri from my friends at PackT.</p>
<h2 id="authors">Authors</h2>
<p><strong>Yousri El Fattah</strong> is the CEO of Causal Computing and an expert in machine intelligence, causal modelling, control systems engineering, and data science.</p>
<p><strong>Reza Bagheri</strong> is a working data scientist at Ipsos. He has written extensively on data science and machine learning, and has spoken at substantial conferences.</p>
<p>Both authors have completed PhDs.</p>
<h2 id="content">Content</h2>
<p>This book is an ambitious and substantial guide to one of the most important areas in modern data science: moving beyond correlation to reason about cause and effect. The book positions Bayesian networks not just as probabilistic modelling tools, but as a practical framework for representing knowledge, handling uncertainty, and supporting causal decision-making. That is a timely focus, as Bayesian networks are increasingly used to combine observational data with domain expertise and to reason about interventions and counterfactuals across fields such as epidemiology, economics, genomics, and environmental science.</p>
<p>The book has amazing breadth. It starts with foundations: probability, Bayes’ theorem, conditional independence, Bayesian networks, and structural causal models. From there, it moves into deeper material including relational database representations of probabilistic models, join tree clustering, belief propagation, and variable elimination. The causal inference chapters then cover key concepts such as Pearl’s do-calculus, back-door and front-door criteria, potential outcomes, counterfactual reasoning, and causal effect identification.</p>
<p>Given the tough topics, the practical nature of the book is welcome. The authors reinforce concepts with worked examples and implementations in R and Python, using packages such as pgmpy, CausalModels, and causallib. Later chapters apply the ideas to real-world-style case studies in economics, epidemiology, and social science, including the Lalonde National Supported Work dataset, smoking cessation and mortality analysis, and the Card and Krueger minimum-wage study. These help lift the book above a purely theoretical treatment and makes it more useful for readers who want to implement causal workflows rather than only understand the mathematics.</p>
<p>This book is not a lightweight introduction. <strong>It assumes readers already have some comfort with probability, statistics, R, Python, and scientific libraries</strong>. That’s already a tough call but at more than 600 pages, it’s also pretty dense reading. I suspect that some readers may find the progression from Bayesian networks into relational algebra and join tree methods quite demanding. I don’t think they should be omitted but readers might find them more specialized than the title initially suggested. This is not an applied <strong>how do I estimate treatment effects?</strong> guide and many will likely find the early and middle chapters heavier than expected.</p>
<p><strong>The best audience is probably data scientists, researchers, and technically confident students who want a serious bridge between graphical models, structural causal models, and applied causal estimation</strong>. For them, the book’s combination of theory, algorithms, and code is excellent. It shows not only what causal inference is trying to achieve, but also how graphical models can clarify assumptions, expose confounding, and guide valid estimation.</p>
<h2 id="summary">Summary</h2>
<p>This book is a rigorous, practical, and unusually broad treatment of the subject. It requires commitment, but for readers who want to understand causal inference at a deeper level and implement it in R or Python, it is a valuable and timely resource.</p>
<p>I liked the book, even though it made for a heavy reading experience. The knowledge-level of the authors is clear.</p>
<p>7 out of 10</p>

      ]]></content:encoded>
    </item>
    
    <item>
      <title>Fabric RTI 101: Providing Fault Tolerance</title>
      <link>https://blog.greglow.com/2026/08/13/fabric-rti-101-providing-fault-tolerance/</link>
      <guid>https://blog.greglow.com/2026/08/13/fabric-rti-101-providing-fault-tolerance/</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 AEST</pubDate>

      <description>Fault tolerance means designing systems that can keep running even when parts of them fail. In real-time data pipelines, something will eventually go wrong — a node might go offline, a network connection could drop, or a consumer might crash. The key is to plan for those failures from the start.
The first principle is to expect failure and aim for graceful degradation rather than total outage. Your system should continue operating, perhaps with reduced functionality or performance, while recovery takes place.
</description>

      <content:encoded><![CDATA[
        
          <img src="https://blog.greglow.com/FabricRTI101.png" alt="cover image" /><br />
        
        <p>Fault tolerance means designing systems that can keep running even when parts of them fail. In real-time data pipelines, something will eventually go wrong — a node might go offline, a network connection could drop, or a consumer might crash. The key is to plan for those failures from the start.</p>
<p>The first principle is to expect failure and aim for graceful degradation rather than total outage. Your system should continue operating, perhaps with reduced functionality or performance, while recovery takes place.</p>
<p><img src="https://greglow.blob.core.windows.net/blog/images/FabricRTI101_10_06_01.png" alt="Providing Fault Tolerance"></p>
<p>One essential mechanism for this is checkpoints and offsets. These allow components such as Eventstreams or consumers to resume from the last successfully processed event after a failure. Instead of replaying all historical data or losing progress, the system simply picks up where it left off.</p>
<p>Another layer of protection comes from partition replicas in your event broker. Systems like Kafka and Azure Event Hubs can replicate partitions across multiple nodes, providing automatic failover if one node goes down. This redundancy ensures message durability and availability even during hardware or service interruptions.</p>
<p>You should also design consumers to be stateless wherever possible. Stateless consumers don’t maintain internal session data or temporary state — they simply process messages as they come in. This makes them much easier to scale horizontally and recover quickly, since a failed instance can be replaced immediately without losing context.</p>
<p>Finally, test fault scenarios before going live. Simulate node failures, network interruptions, and restart cycles to confirm that your pipeline resumes correctly and maintains data integrity. Testing under failure conditions is the only way to verify true resilience.</p>
<p>Providing fault tolerance means planning for failure, not reacting to it. By using checkpoints, replicas, stateless components, and thorough testing, you can ensure your real-time system continues to operate reliably — even when individual pieces break.</p>
<h2 id="learn-more-about-fabric-rti">Learn more about Fabric RTI</h2>
<p>If you really want to learn about RTI right now, we have an online on-demand course that you can enrol in, right now. You’ll find it at 





  <a href="https://sqldownunder.com/courses/rti">Mastering Microsoft Fabric Real-Time Intelligence</a>

</p>

      ]]></content:encoded>
    </item>
    
    <item>
      <title>SQL: Odd TRY_CAST and TRY_CONVERT Behavior</title>
      <link>https://blog.greglow.com/2026/08/12/sql-odd-try_cast-and-try_convert-behavior/</link>
      <guid>https://blog.greglow.com/2026/08/12/sql-odd-try_cast-and-try_convert-behavior/</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 AEST</pubDate>

      <description>Here’s a quick T-SQL test for you.
Without looking below to see the answer first, try to guess what each of these statements will produce as output:
SELECT TRY_CAST('' AS int); SELECT TRY_CAST(' ' AS int); SELECT TRY_CAST('' AS date); SELECT TRY_CAST('' AS decimal(18, 2)); SELECT TRY_CONVERT(date, '', 103); And to slightly distract you from checking out the answers yet, here is another wise-looking owl who is thinking about the answers, and warning you not to look further down the page yet:
</description>

      <content:encoded><![CDATA[
        
          <img src="https://blog.greglow.com/Owl1.png" alt="cover image" /><br />
        
        <p>Here’s a quick T-SQL test for you.</p>
<p>Without looking below to see the answer first, try to guess what each of these statements will produce as output:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-sql" data-lang="sql"><span style="display:flex;"><span><span style="color:#66d9ef">SELECT</span> TRY_CAST(<span style="color:#e6db74">''</span> <span style="color:#66d9ef">AS</span> int);
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">SELECT</span> TRY_CAST(<span style="color:#e6db74">'    '</span> <span style="color:#66d9ef">AS</span> int);
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">SELECT</span> TRY_CAST(<span style="color:#e6db74">''</span> <span style="color:#66d9ef">AS</span> date);
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">SELECT</span> TRY_CAST(<span style="color:#e6db74">''</span> <span style="color:#66d9ef">AS</span> decimal(<span style="color:#ae81ff">18</span>, <span style="color:#ae81ff">2</span>));
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">SELECT</span> TRY_CONVERT(date, <span style="color:#e6db74">''</span>, <span style="color:#ae81ff">103</span>);
</span></span></code></pre></div><p>And to slightly distract you from checking out the answers yet, here is another wise-looking owl who is thinking about the answers, and warning you not to look further down the page yet:</p>
<p><img src="https://greglow.blob.core.windows.net/blog/images/Owl2.png" alt=""></p>
<p>Anyway, here’s what happens when you run this T-SQL in SQL Server:</p>
<p><img src="https://greglow.blob.core.windows.net/blog/images/TryCast_Image2.png" alt=""></p>
<p>Surprised? I’d have to say that I was. Now as my buddy Adam Machanic pointed out, it’s not the fault of TRY_CAST and TRY_CONVERT because they just TRY to do a CAST and a CONVERT. And it’s the original functions that have the bizarre behavior, not the TRY versions of them.</p>
<p>Can’t say that I love this because it means that I can’t use these functions for their purpose, except for <strong>decimal</strong>. So that then left me wondering which types had this behavior.</p>
<p>Let’s find out! I executed the following:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-sql" data-lang="sql"><span style="display:flex;"><span>USE tempdb;
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">GO</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">SET</span> NOCOUNT <span style="color:#66d9ef">ON</span>;
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">DECLARE</span> <span style="color:#f92672">@</span>TypeName sysname;
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">DECLARE</span> <span style="color:#f92672">@</span><span style="color:#66d9ef">SQL</span> nvarchar(<span style="color:#66d9ef">max</span>);
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">DECLARE</span> <span style="color:#f92672">@</span>Outcomes <span style="color:#66d9ef">TABLE</span>
</span></span><span style="display:flex;"><span>(
</span></span><span style="display:flex;"><span>    OutcomeID int <span style="color:#66d9ef">IDENTITY</span>(<span style="color:#ae81ff">1</span>,<span style="color:#ae81ff">1</span>) <span style="color:#66d9ef">PRIMARY</span> <span style="color:#66d9ef">KEY</span>,
</span></span><span style="display:flex;"><span>    TypeName sysname,
</span></span><span style="display:flex;"><span>    ReturnedValue sql_variant
</span></span><span style="display:flex;"><span>);
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">DECLARE</span> TypeList <span style="color:#66d9ef">CURSOR</span> FAST_FORWARD READ_ONLY
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">FOR</span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">SELECt</span> typ.[name] <span style="color:#66d9ef">AS</span> TypeName
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">FROM</span> sys.types <span style="color:#66d9ef">AS</span> typ
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">WHERE</span> typ.system_type_id <span style="color:#f92672">=</span> typ.user_type_id
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">AND</span> typ.[name] <span style="color:#66d9ef">NOT</span> <span style="color:#66d9ef">IN</span> (N<span style="color:#e6db74">'image'</span>, N<span style="color:#e6db74">'json'</span>, N<span style="color:#e6db74">'text'</span>, N<span style="color:#e6db74">'ntext'</span>, N<span style="color:#e6db74">'timestamp'</span>, N<span style="color:#e6db74">'xml'</span>)
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">ORDER</span> <span style="color:#66d9ef">BY</span> TypeName;
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">OPEN</span> TypeList;
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">FETCH</span> <span style="color:#66d9ef">NEXT</span> <span style="color:#66d9ef">FROM</span> TypeList <span style="color:#66d9ef">INTO</span> <span style="color:#f92672">@</span>TypeName;
</span></span><span style="display:flex;"><span>WHILE <span style="color:#f92672">@@</span>FETCH_STATUS <span style="color:#f92672">=</span> <span style="color:#ae81ff">0</span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">BEGIN</span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">SET</span> <span style="color:#f92672">@</span><span style="color:#66d9ef">SQL</span> <span style="color:#f92672">=</span> N<span style="color:#e6db74">'SELECT '''</span> <span style="color:#f92672">+</span> <span style="color:#f92672">@</span>TypeName <span style="color:#f92672">+</span> N<span style="color:#e6db74">''', TRY_CAST('''' AS '</span> <span style="color:#f92672">+</span> <span style="color:#f92672">@</span>TypeName <span style="color:#f92672">+</span> N<span style="color:#e6db74">');'</span>;
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">INSERT</span> <span style="color:#f92672">@</span>Outcomes (TypeName, ReturnedValue)
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">EXEC</span>(<span style="color:#f92672">@</span><span style="color:#66d9ef">SQL</span>);
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>    <span style="color:#66d9ef">FETCH</span> <span style="color:#66d9ef">NEXT</span> <span style="color:#66d9ef">FROM</span> TypeList <span style="color:#66d9ef">INTO</span> <span style="color:#f92672">@</span>TypeName;
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">END</span>;
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">CLOSE</span> TypeList;
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">DEALLOCATE</span> TypeList;
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">SELECT</span> <span style="color:#f92672">*</span> <span style="color:#66d9ef">FROM</span> <span style="color:#f92672">@</span>Outcomes <span style="color:#66d9ef">ORDER</span> <span style="color:#66d9ef">BY</span> TypeName;
</span></span></code></pre></div><p>Note that I excluded old data types, and others that can’t cast to sql_variant anyway. (Mind you, no idea why XML and JSON can’t be cast to sql_variant). And here’s the outcome:</p>
<p><img src="https://greglow.blob.core.windows.net/blog/images/TryCast_Image4.png" alt=""></p>
<p>So, you’ve been warned.</p>

      ]]></content:encoded>
    </item>
    
    <item>
      <title>Fabric RTI 101: Handling Retries</title>
      <link>https://blog.greglow.com/2026/08/11/fabric-rti-101-handling-retries/</link>
      <guid>https://blog.greglow.com/2026/08/11/fabric-rti-101-handling-retries/</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 AEST</pubDate>

      <description>In any real-time system, we can’t assume that every event will be delivered successfully on the first attempt. Network interruptions, temporary service outages, or throttling limits can all cause event delivery failures.
That’s why retry logic is a fundamental part of reliable streaming architecture.
The most common approach is to use exponential backoff — meaning that the system waits progressively longer between retries. This avoids overwhelming the destination service during outages. For example, retries might occur after 1 second, then 2 seconds, then 4, and so on, up to a maximum delay.
</description>

      <content:encoded><![CDATA[
        
          <img src="https://blog.greglow.com/FabricRTI101.png" alt="cover image" /><br />
        
        <p>In any real-time system, we can’t assume that every event will be delivered successfully on the first attempt. Network interruptions, temporary service outages, or throttling limits can all cause event delivery failures.</p>
<p>That’s why retry logic is a fundamental part of reliable streaming architecture.</p>
<p>The most common approach is to use exponential backoff — meaning that the system waits progressively longer between retries. This avoids overwhelming the destination service during outages. For example, retries might occur after 1 second, then 2 seconds, then 4, and so on, up to a maximum delay.</p>
<p>However, retries can introduce another challenge: duplicate events. If a message is retried and the previous attempt eventually succeeds, both versions could be processed.</p>
<p><img src="https://greglow.blob.core.windows.net/blog/images/FabricRTI101_10_05_01.png" alt="Handling Retries"></p>
<p>To handle this safely, design your system to be idempotent — in other words, processing the same event twice should produce the same outcome. This might involve checking event IDs, timestamps, or sequence numbers to ensure no duplication in the final results.</p>
<p>Another best practice is to implement poison message queues. These are special queues used to isolate problematic events that consistently fail processing. By diverting these messages away from the main pipeline, you prevent a few bad events from blocking or slowing down the entire stream.</p>
<p>Finally, make sure to log all retries — both successful and failed ones. This helps with monitoring, troubleshooting, and root cause analysis. If a service starts to show an increasing retry rate, it’s often an early sign of underlying performance or connectivity problems.</p>
<p>Handling retries effectively is about resilience and control. You accept that failures will happen, but you ensure that when they do, the system recovers gracefully, avoids duplication, and provides the observability you need to fix the issue quickly.</p>
<h2 id="learn-more-about-fabric-rti">Learn more about Fabric RTI</h2>
<p>If you really want to learn about RTI right now, we have an online on-demand course that you can enrol in, right now. You’ll find it at 





  <a href="https://sqldownunder.com/courses/rti">Mastering Microsoft Fabric Real-Time Intelligence</a>

</p>

      ]]></content:encoded>
    </item>
    
    <item>
      <title>Fix: Failed to update the database because the database is read-only</title>
      <link>https://blog.greglow.com/2026/08/10/fix-failed-to-update-the-database-because-the-database-is-read-only/</link>
      <guid>https://blog.greglow.com/2026/08/10/fix-failed-to-update-the-database-because-the-database-is-read-only/</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 AEST</pubDate>

      <description>Had a client today asking about this error message. They were working away on a machine and suddenly they got the message Failed to update the database because the database is read-only.
The user hadn’t changed anything that they were aware of. Based on the user’s permissions (ie: what they could see), everything in SSMS looked normal. When they checked the sys.databases view, the database showed MULTI_USER. There was enough disk space. Folder permissions had not changed. The user was puzzled. The issue was caused by the database being part of an availability group, and the AG had failed over. So suddenly, the database the user was connected to, was now a read-only replica, not the primary replica. That’s why the database said it was read-only.
</description>

      <content:encoded><![CDATA[
        
          <img src="https://blog.greglow.com/Fix.png" alt="cover image" /><br />
        
        <p>Had a client today asking about this error message. They were working away on a machine and suddenly they got the message <strong>Failed to update the database because the database is read-only</strong>.</p>
<ul>
<li>The user hadn’t changed anything that they were aware of.</li>
<li>Based on the user’s permissions (ie: what they could see), everything in SSMS looked normal.</li>
<li>When they checked the sys.databases view, the database showed MULTI_USER.</li>
<li>There was enough disk space.</li>
<li>Folder permissions had not changed.</li>
<li>The user was puzzled.</li>
</ul>
<p>The issue was caused by the database being part of an availability group, and the AG had failed over. So suddenly, the database the user was connected to, was now a read-only replica, not the primary replica. That’s why the database said it was read-only.</p>
<p>What the user should have done was to have connected to the AG listener in the first place, not to the server name. Then when failover occurs, the listener would follow the primary server.</p>
<p>I think this error message is confusing. I really wish it gave you a hint about what’s going on.</p>

      ]]></content:encoded>
    </item>
    
    <item>
      <title>Fabric RTI 101: Handling Backpressure</title>
      <link>https://blog.greglow.com/2026/08/09/fabric-rti-101-handling-backpressure/</link>
      <guid>https://blog.greglow.com/2026/08/09/fabric-rti-101-handling-backpressure/</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 AEST</pubDate>

      <description>In any real-time data system, there’s a point where the incoming event rate can exceed what the system can process. This condition is known as backpressure.
Backpressure can occur for a few reasons — a sudden spike in data volume, slow or overloaded consumers, or limited throughput in one part of the pipeline. If it’s not handled properly, it can cascade through the system, eventually causing delays or even a complete stall in event processing.
</description>

      <content:encoded><![CDATA[
        
          <img src="https://blog.greglow.com/FabricRTI101.png" alt="cover image" /><br />
        
        <p>In any real-time data system, there’s a point where the incoming event rate can exceed what the system can process. This condition is known as backpressure.</p>
<p>Backpressure can occur for a few reasons — a sudden spike in data volume, slow or overloaded consumers, or limited throughput in one part of the pipeline. If it’s not handled properly, it can cascade through the system, eventually causing delays or even a complete stall in event processing.</p>
<p>There are several strategies to handle backpressure effectively.</p>
<p><img src="https://greglow.blob.core.windows.net/blog/images/FabricRTI101_10_04_01.png" alt="Handling Backpressure"></p>
<p>The first is buffering. By introducing temporary storage — for example, using queues or durable event hubs — you can smooth out short bursts of high activity. Buffers don’t eliminate the problem, but they give downstream components time to catch up.</p>
<p>The next approach is to scale out consumers or processing nodes. When event volume increases, adding parallel processing capacity allows multiple consumers to share the workload. Partitioned event streams make this approach more effective, since each partition can be processed independently.</p>
<p>For longer-term or persistent overload conditions, you may need to prioritize certain events. In some cases, it’s acceptable to drop low-priority or redundant events to maintain system responsiveness. For instance, telemetry systems might sample events instead of processing every single one when volume peaks.</p>
<p>Monitoring plays a key role here — metrics such as queue length, consumer lag, and end-to-end latency can help identify when backpressure is forming.</p>
<p>The overall goal is to prevent a total stall. A well-designed system should degrade gracefully under heavy load — slowing down processing or dropping nonessential data instead of failing outright.</p>
<p>Handling backpressure is about maintaining stability and responsiveness when your real-time pipeline is under stress.</p>
<h2 id="learn-more-about-fabric-rti">Learn more about Fabric RTI</h2>
<p>If you really want to learn about RTI right now, we have an online on-demand course that you can enrol in, right now. You’ll find it at 





  <a href="https://sqldownunder.com/courses/rti">Mastering Microsoft Fabric Real-Time Intelligence</a>

</p>

      ]]></content:encoded>
    </item>
    
    <item>
      <title>Book Review: Python for Algorithmic Trading Cookbook</title>
      <link>https://blog.greglow.com/2026/08/08/book-review-python-for-algorithmic-trading-cookbook/</link>
      <guid>https://blog.greglow.com/2026/08/08/book-review-python-for-algorithmic-trading-cookbook/</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 AEST</pubDate>

      <description>I recently received a review copy of Python for Algorithmic Trading Cookbook (2nd Edition): Recipes for designing, building, and deploying algorithmic trading strategies with Python by Jason Strimpel from my friends at PackT.
Author Jason Strimpel is the founder of PyQuant News, co-founder of Quant Science, and Managing Director of Global AI and Advanced Analytics at a top-tier consulting firm.
Content This book is an ambitious, highly practical guide to building the complete research-to-execution workflow for systematic trading. It treats algorithmic trading as an engineering discipline involving data acquisition, storage, analysis, backtesting, risk measurement, execution, and deployment. I was surprised by the breadth of modern tools that are discussed: 68 recipes, 51 Jupyter notebooks, 17 modular trading applications, and four GPU-focused scripts.
</description>

      <content:encoded><![CDATA[
        
          <img src="https://blog.greglow.com/PythonForAlgorithmicTradingCookbook_BookCover.png" alt="cover image" /><br />
        
        <p>I recently received a review copy of 





  <a href="https://www.packtpub.com/en-au/product/python-for-algorithmic-trading-cookbook-9781806662029">Python for Algorithmic Trading Cookbook (2nd Edition): Recipes for designing, building, and deploying algorithmic trading strategies with Python</a>

 by Jason Strimpel from my friends at PackT.</p>
<h2 id="author">Author</h2>
<p><strong>Jason Strimpel</strong> is the founder of PyQuant News, co-founder of Quant Science, and Managing
Director of Global AI and Advanced Analytics at a top-tier consulting firm.</p>
<h2 id="content">Content</h2>
<p>This book is an ambitious, highly practical guide to building the complete research-to-execution workflow for systematic trading. It treats algorithmic trading as an engineering discipline involving data acquisition, storage, analysis, backtesting, risk measurement, execution, and deployment. I was surprised by the breadth of modern tools that are discussed: 68 recipes, 51 Jupyter notebooks, 17 modular trading applications, and four GPU-focused scripts.</p>
<p>I like cookbook-style books. Often, I don’t want to learn every aspect of a concept, but just want to see examples of how to use it. The recipes in the book mostly begin with the required setup, move through implementation steps in order, explain how the code works, and then suggest extensions. This structure makes the book suitable both for sequential study and for later use as a reference.</p>
<p>Jason’s explanations are direct and based on real tasks rather than abstract demonstrations. For example, he doesn’t just teach basic DataFrame operations; instead, the examples use pandas, Polars, DuckDB, Parquet, and ArcticDB to process and manage substantial financial datasets. The ArcticDB material is especially valuable because it addresses reproducibility and look-ahead bias through versioned, point-in-time data.</p>
<p>I also like the way that the examples connect across chapters. Data prepared earlier is reused in later research, backtesting, and deployment exercises, so the book feels more coherent than a collection of unrelated snippets. It’s always hard to know whether to take this approach or not. It lets you build something more substantial, but can make it harder for someone who wants to just dive in later. It works best for someone reading the whole book. The used of procedural notebooks keeps the early material approachable, while the modular application built in later chapters introduces separation of clients, wrappers, contracts, orders, and utilities without becoming too architectural.</p>
<p>The best part of the book is its end-to-end scope. It progresses from sourcing equities, futures, options, and factor data through visualisation, alpha-factor construction, vectorised and event-driven backtesting, factor evaluation, portfolio analytics, and live brokerage integration. The backtesting chapters also go beyond reporting attractive returns. Walk-forward testing is used to investigate overfitting, while Zipline Reloaded examples incorporate commissions and slippage. Alphalens and Pyfolio recipes examine information coefficients, turnover, drawdowns, exposures, transaction costs, and trade-level performance. This emphasis on robustness distinguishes the book from many introductory trading texts. That’s amazing detail for this type of book.</p>
<p>The later chapters seem quite current. A substantial section explores AI-assisted research with LangChain, LlamaIndex, retrieval-augmented generation, and multi-agent workflows. Importantly, Jason presents these systems as tools for accelerating research rather than replacing human judgement. The chapters on Interactive Brokers build reusable application components for contracts, orders, streaming data, positions, portfolio profit and loss, and paper or live deployment. This gives readers a credible path beyond notebooks and into operational systems.</p>
<p>Are there any limitations? Sure. The breadth of topics means that several sophisticated ones receive recipe-sized treatments rather than deep theoretical development. This is not a university textbook on the topic. So readers seeking rigorous derivations of factor models, portfolio optimisation, market microstructure, or statistical testing will need additional reading material. I could also see some potential installation issues with the large technology stack that’s used, particularly because several libraries evolve quickly and some examples require API keys, premium data, an Interactive Brokers account, or specialised configuration. The final GPU chapter is impressive, but if you want to follow its largest examples, you’ll need appropriate NVIDIA hardware and considerable memory, although smaller datasets can be substituted.</p>
<p>This is clearly not an introductory-level first Python book. <strong>Readers should already understand basic syntax, pandas, NumPy, and common financial terminology</strong>. The real value in the book is in providing templates for disciplined experimentation.</p>
<h2 id="summary">Summary</h2>
<p>I didn’t ever get to read the first edition of this book, but this second edition is an excellent resource for developers, quantitatively-minded traders, and investors who want to understand the machinery surrounding a strategy and not merely its entry and exit rules. It combines breadth, practical code, realistic cautions, and a strong progression from research data to deployed trading applications. While I could see some challenges in its dependencies (as they are moving targets), this has to be one of the more comprehensive hands-on guides to the modern Python algorithmic-trading ecosystem.</p>
<p>9 out of 10</p>

      ]]></content:encoded>
    </item>
    
    <item>
      <title>Fabric RTI 101: Designing for Reliability and Resilience</title>
      <link>https://blog.greglow.com/2026/08/07/fabric-rti-101-designing-for-reliability-and-resilience/</link>
      <guid>https://blog.greglow.com/2026/08/07/fabric-rti-101-designing-for-reliability-and-resilience/</guid>
      <pubDate>Fri, 07 Aug 2026 00:00:00 AEST</pubDate>

      <description>In real-time systems, reliability and resilience are essential. Unlike batch workloads, where you can retry a job later, streaming pipelines are continuous — they must keep running even when parts of the system fail.
The first design principle is to ensure that streams keep flowing.
That often means building redundancy into both your data sources and destinations. For example, you might configure multiple input connections or have failover routes so that data can still be processed if one stream or endpoint goes offline.
</description>

      <content:encoded><![CDATA[
        
          <img src="https://blog.greglow.com/FabricRTI101.png" alt="cover image" /><br />
        
        <p>In real-time systems, reliability and resilience are essential. Unlike batch workloads, where you can retry a job later, streaming pipelines are continuous — they must keep running even when parts of the system fail.</p>
<p>The first design principle is to ensure that streams keep flowing.</p>
<p>That often means building redundancy into both your data sources and destinations. For example, you might configure multiple input connections or have failover routes so that data can still be processed if one stream or endpoint goes offline.</p>
<p><img src="https://greglow.blob.core.windows.net/blog/images/FabricRTI101_10_03_01.png" alt="Designing for Reliability and Resilience"></p>
<p>You should also design for failover and recovery. Every major component — such as an Eventstream, KQL Database, or downstream output — should be capable of restarting and catching up automatically after an interruption. This might involve using checkpoints or replay mechanisms to resume from the last known offset rather than losing messages.</p>
<p>Another key concept is message durability. Event brokers like Kafka, Azure Event Hubs, or Service Bus can be configured to persist events to durable storage. This ensures that even if a consumer or service crashes, the events can be replayed later. The level of durability affects performance and cost, so it’s important to configure it appropriately for your reliability requirements.</p>
<p>Ongoing monitoring is also part of resilience. Two key indicators are lag — how far behind the consumers are from the producers — and throughput, which measures how quickly events are processed. Tracking these metrics helps detect performance degradation early, before it escalates into downtime.</p>
<p>Designing for reliability means expecting failure and planning how the system will recover. Resilient real-time architectures rely on redundancy, durability, and visibility to ensure that even when something breaks, the data keeps flowing and the system recovers gracefully.</p>
<h2 id="learn-more-about-fabric-rti">Learn more about Fabric RTI</h2>
<p>If you really want to learn about RTI right now, we have an online on-demand course that you can enrol in, right now. You’ll find it at 





  <a href="https://sqldownunder.com/courses/rti">Mastering Microsoft Fabric Real-Time Intelligence</a>

</p>

      ]]></content:encoded>
    </item>
    
  </channel>
</rss>
