Enable javascript in your browser for better experience. Need to know to enable it? Go here.

AI-Ready Data: The anthology

Part 2

In part 1 of the AIRD anthology, our focus was on foundational prerequisites for an AI-ready platform. This included ensuring the data fueling the platform was high quality and governed securely so it could be trusted; using metadata and cataloguing to improve efficiency across a data-first organization; and establishing a semantic business layer so AI systems could identify and use the right dataset for each task.

 

Once a solid foundation is built, more sophisticated components and topics can be introduced. This is the overarching focus of Part two. 



A to B accessibility

 

As this series outlines, access to data is a key requirement for AI systems. There are two common ways to enable accessibility: APIs and SQL. This battle of the TLAs (three-letter acronyms) is really a question of governance versus flexibility. Accessing data via APIs typically requires more setup and results in higher latency, but it is a more secure approach as it leverages tools to access the data via defined endpoints. 

 

Data access via SQL means the AI needs to be able to understand relationships across the data; it requires less setup as the AI can directly access the underlying data, but it can introduce additional risks like hallucinations and SQL injection.

 

A well-funded and reliable public transit system offers a variety of means of transportation and routes. Users of the system can plan their trips based on the various options available as long as a nearby stop or station exists. While this means of transit may not take them door-to-door, it is a safe and predictable (perhaps rigid) way to get from A to B, just as APIs are pre-defined, safe routes with which AI can interact with data.

 

Meanwhile, driving your own vehicle means you have the flexibility to move between the exact locations on your own schedule. Your own route, including beyond traditional roadways, is also an option. However, the risk of getting lost or unpredictable timing is significantly higher. The flexible, custom paths made available by SQL are like driving, and for this reason, guardrails (i.e., a driver’s license or permissions on the types of queries which can be executed) are especially critical.

 

 

The case of unstructured knowledge management

 

One of the major benefits of AI is its ability to ingest, process and utilize a wide variety and large volume of data types. Most people are familiar with structured data: data which lives in tables or spreadsheets with a fairly fixed set of headers (columns). However, AI can also process high volumes of unstructured data. 

 

This includes images, videos and documents; basically anything lacking fixed schemas or predefined fields. AI does so by scanning and parsing the inputs for plain text as well as metadata, breaking these inputs down into smaller units and then creating vectors which it stores and accesses to perform its operations. One of the challenges around this are ensuring the vectors are kept up-to-date, even when the source data may change; and it may also be tricky to ensure original source permissions are followed by the AI system. In the breakdown to smaller units, there is a high risk of separating the context from a given datapoint.

 

Detective Caitlyn is working on a complex case and doesn’t have an objective source of truth, such as incriminating surveillance footage, clear fingerprints that she can easily process through a database or completely trustworthy interviewee statements. Instead, Detective Caitlyn has a messy crime scene as her starting point. She needs to process the crime scene by focusing on the important parts, like bagging evidence and dusting for fingerprints. Detective Caitlyn then needs to work with various parties to make sense of the evidence: crime scene unit investigators, fingerprint databases, witnesses. 

 

This may bring us to a wall of evidence and notes connected by red string. Individually, each piece does not tell the whole story. In a data ecosystem, that red string is metadata linking disparate data sources. Connecting the pieces via the red string helps to give context to the components, and build the narrative. However, not all the evidence or leads will be pertinent to closing the case, and only the relevant information describing what happened should be presented to a prosecutor.  

 

 

Healthmaxxing lineage and observability

 

Lineage and observability are two important concepts in ensuring an AI system is behaving as intended. They are often sub-topics under the umbrella of governance, but deserve to have their own anecdotes. Lineage answers “How did we get here?”, and observability asks, “How are we doing right now?”. A well-built observability system enables a comprehensive lineage system by providing it with information. Meanwhile, the lineage system can be leveraged to explain what’s being captured by the observability mechanisms.

 

Sophie is interested in leveraging her smart watch and paired app (observability) to support her healthy lifestyle. If she were to just start tracking now without any medical history (lineage), the monitoring systems may still be useful, but she has no real baseline to go off of. Her first instance of each scenario in using the technology would be the comparison point. If her heart rate is above average to start with, she might not be aware of a family history of potential high blood pressure. 

 

However, if Sophie had a well-tracked medical record, including blood pressure tests, physical exams and her family history, she may be able to correlate certain events in her tracking to probable underlying causes. The past information on its own may also still be useful to watch out for certain conditions, but without fresh data, the flags may go unnoticed. The health of the whole system, Sophie, is vastly improved by having a well-structured historical record as well as a performant real-time monitoring mechanism. 



Vector/RAG readiness audit

 

This is where it all comes together. Many organizations assume that having high volumes and unfettered access to data is sufficient for vector and retrieval-augmented generation (RAG) applications. More is not always better. Often, documents that are optimized for human consumption are not optimized for machine consumption. The data needs to be in a state for the organization’s internal knowledge to connect to large language models (LLMs).

 

Josh is an auditor who has just arrived on-site with his newest client. Josh has a process that he follows on the job, and it has worked well across a variety of companies and industries. However, when Josh arrives at this new client, the information he is passed is a series of receipts from different transactions and vendors, random pieces of paper with calculations scrawled on them and conflicting invoices. Somehow, this company appears to be successful, but their processes for managing the books are not conducive to an efficient (or effective) audit.

 

An audit could be performed much more efficiently if the company had all their information digitized in a consistent format (i.e., spreadsheets) with complete dates, balances and memos. Better yet, a bookkeeping software with a record of financial operations would make analyzing the company’s activity far more efficient, as well as increase confidence in the calculations to be performed (and those performed historically).

 

Just like Josh, AI needs to work with data that is hygienic (common formats, minimal missing information) as well as includes complete context (rich metadata and tagging). This will enable the AI to efficiently use tools for parsing, embedding, storage and retrieval. It also enables a more secure setup with permissions management.

 

 

Real-time capabilities with a side of fries

 

“Walk before you run” is apt for developing data platforms. Of course, even walking requires a solid foundation (see all of the above). Think of this as the batch AI data platform. Once you have the hang of that, you can consider picking up the pace to enable real-time capabilities.

 

Drive-in restaurants were popular in the 1920s. These establishments were similar to diners, where patrons placed an order and then waited for their food. However, instead of going into the restaurant, customers placed the order and waited in their parked cars for a server (often on rollerskates) to bring them their food. The whole session is tied to a single stall (wherever the car is parked), and when the customer is finished, they clear the stall for the next set of processing to fill it.

 

This process started to increase its pace in the 1930s and the decades that followed. In the 1970s, the drive-up window became popular. This shifted the process from being a parked, vehicle-off experience to a queue of vehicles streaming in and out of a lane to pick up food which most likely began being prepared before the order was placed. The whole system is low-latency with minimal hand-off required. However, if one order or vehicle stalls up the line, the latency of the entire system spikes.

 

Similarly, data for real-time AI applications must exist in a stream with accurate, precise and standard-format timestamps, so if something is processed out-of-order, the sessions can be associated with specific users or operations and the state of the dynamic system can be maintained. Additionally, in order to make use of the real-time capabilities of the system, the data must immediately update the operational memory of the engine. Failure to do so risks performing operations with stale inputs. Furthermore, as consumers of a real-time system are likely expecting at least near real-time results, the system must be able to serve the data with minimal latency. 

 

 

You need brakes to go fast

 

As the technologies which enable AI and its applications get advanced, focusing on the foundations of the data being fed into these systems becomes increasingly important. It is critical that the acceleration and amplification by AI are developed on a solid foundation. Autonomous multi-agent systems, speech-to-speech systems and chain of reasoning are all powerful and complex uses of AI, which are only as good (technically and ethically) as the data they are fed.

 

Each anecdote in this article highlights key principles of responsible (and necessary) data preparation for AI systems. One good test for whether or not an organization’s data is truly AI-ready is to answer the question: can a new employee understand, find, trust and use the data without asking five people for help? If the answer is “yes”, the organization has some important foundations for AI-ready data already in place. If the answer is “no”, it may be good to revisit some of the anecdotes in this article to determine where improvements must be made.

Disclaimer: The statements and opinions expressed in this article are those of the author(s) and do not necessarily reflect the positions of Thoughtworks.

Explore a snapshot of today's tech landscape