The Machine Gaze: How Trevor Paglen Maps the Erosion of Visual Reality
Artist Trevor Paglen examines how computer vision and generative media have transformed images into operational tools of power and control.

In the sprawling, often unseen workshops of enterprise technology, a quiet revolution is underway.
It’s a space where data, once an ethereal stream of information, is increasingly becoming a tangible asset – a raw material, a refined component, even a finished product.
This isn’t just about managing information anymore; it’s about engineering it, meticulously shaping it for a singular, overarching purpose: the insatiable demands of artificial intelligence.
For too long, the discussions around this foundational work have been cloaked in the more palatable term data science.
But beneath the academic veneer of algorithms and models lies the gritty reality of data mechanics—a discipline rooted in the practical, hands-on work of forging, refining, and delivering data.
It’s a field that, despite being well over half a century old with the advent of the first database systems, is now experiencing an unprecedented surge of innovation, driven by the undeniable imperative of AI.
The central dogma echoing through these workshops is starkly simple: “Okay, you’ve got data, but does your data work well for AI?”
As any seasoned technologist knows, AI is only as smart as the data it’s fed.
The old adage, “garbage in, garbage out,” has never been more pertinent.
This isn’t a passing trend; it’s the bedrock upon which the entire AI edifice is being built, demanding a fundamental re-evaluation of how data is prepared, trusted, and deployed across every facet of IT, from DevOps to ERP systems.
Consider the strategic maneuvers of companies like NetApp, a name synonymous with intelligent data infrastructure.
Their journey through the data mechanics landscape offers a telling microcosm of the industry’s evolving focus.
Once acquiring Data Mechanics in 2021 to capitalize on Apache Spark’s big data prowess, NetApp has since divested some of these assets, refocusing its formidable storage competencies squarely on the burgeoning AI industry.
This isn’t merely shedding dead weight; it’s a strategic pivot towards providing crucial data engineering resources for AI, exemplified by partnerships with giants like Nvidia for its AI Data Platform reference design via NetApp AIPod, and Intel for the compact AIPod Mini, designed to streamline AI inferencing.
These moves underscore a critical insight: the value now lies not just in storing data, but in ensuring it is perfectly sculpted for AI’s voracious appetite.
Operating as an independent business unit of Hitachi, Pentaho has coined the term data fitness for the age of AI, a concept that perfectly encapsulates the industry’s current obsession.
Pentaho is actively enhancing its Data Catalog, transforming it into an indispensable tool for data operations management.
Imagine a master inventory for a vast, complex factory floor: this catalog helps data scientists and developers not only locate their data but also understand its provenance, monitor its quality, classify its content, and ensure its compliance.
As Kunju Kashalikar, product management executive at Pentaho, aptly puts it, “The need for strong data foundations has never been higher… They want to improve the organization of data for operations and AI.
They need better visibility into the ‘what and where’ of data’s lifecycle for quality, trust and regulations.”
The goal is clear: automate the heavy lifting of data wrangling, ensuring trust in a bewildering mix of custom, licensed, anonymized, and legacy datasets.
The very language of this domain is steeped in industrial metaphors.
The data pipeline is perhaps the most pervasive, conjuring images of raw materials flowing through a series of filters, transformations, and joins, ultimately arriving at its designated endpoint—an application, another service, or crucially, an AI ingestion point.
Technology vendors eagerly lay claim to “end-to-end data pipelines,” a testament to their ambition to span the entire journey of data.
Databricks, a key player in data platforms, has even open-sourced its core declarative extract, transform, and load (ETL) framework as Apache Spark Declarative Pipelines.
Matei Zaharia, Databricks CTO, emphasizes its role in tackling “one of the biggest challenges in data engineering,” making it easier to build reliable, scalable data pipelines for both batch and streaming workloads.
This open-sourcing democratizes access to “battle-tested” frameworks, allowing engineers to focus on business value rather than infrastructure complexities, as Jian Zhou, a senior engineering manager at Navy Federal Credit Union, enthusiastically confirms.
Yet, as the AI frontier expands, particularly with the rise of large language models (LLMs), the very nature of data mechanics is evolving beyond traditional pipelines.
Ken Exner, chief product officer at Elastic, offers a provocative insight: the real challenge for preparing data for LLMs isn’t about formatting; it’s about retrieval and relevance.
LLMs, he argues, are already adept at interpreting raw, unstructured data.
The critical hurdle is getting the right private data to the LLM at the right time, while preserving context, respecting permissions, and enforcing enterprise-grade security.
This goes far beyond the capabilities of conventional ETL, demanding systems that can bridge structured and unstructured data, understand real-time context, and make internal data discoverable and usable, not just clean.
This seismic shift ultimately brings us to the concept of the data product.
Data, once an intangible byproduct, is now a carefully engineered component, a functional entity on the factory floor, as tangible as a server or an application.
The relentless march of AI has transformed data mechanics from a niche discipline into the linchpin of modern enterprise.
The workshops are bustling, the tools are evolving, and the skilled “data mechanics” are painstakingly greasing the wheels of innovation, ensuring that the mountains of often-siloed private data can finally unlock the true promise of generative AI.
This isn’t just about making data available; it’s about making it work.
Artist Trevor Paglen examines how computer vision and generative media have transformed images into operational tools of power and control.
Abhishek Saxena’s work at Sentient targets the gap where open-source AI keeps losing: not capability, but economics—and his answer is infrastructure that automatically pays builders, maintainers, and evaluators every time their artifact is used, enforced by smart contracts rather than legal goodwill. By combining cryptographic fingerprinting, on-chain attribution, and grant funding with no equity attached, Sentient is building the coordination layer that would make open-source development financially rational enough to compete with a corporate salary.
As AI-generated content floods the digital landscape, Substack is empowering readers to verify the human origins of the newsletters they consume.