The Human Element in an Automated World: How Intelligence Secured $7.9M to Solve AI’s Biggest Bottleneck

By Russell Brandom
Published in TechCrunch


Executive Overview

The artificial intelligence boom has long promised a future where machines can generate anything—from code and complex data structures to immersive video games and breathtaking digital art—at the push of a button. Yet, as frontier labs push the boundaries of large-scale media generation, they have run into an unexpected, stubborn barrier: automated models can easily produce functional outputs, but they struggle immensely to produce outputs that are actually good.

Machine learning algorithms can effortlessly compile game mechanics or render high-resolution layouts, but they lack the intrinsic spark of human taste, intuition, and enjoyment. When co-founder Grace Li and a small collective of college friends set out to build an AI-powered game engine shortly before their graduation in 2025, they realized this fundamental shortcoming firsthand. Their early models could construct fully operational games, but none of them were fun.

This realization prompted a pivotal question that would alter the trajectory of their careers: How can an AI system accurately determine whether a creative output will resonate with a human audience?

The answer, they concluded, was that there is simply no substitute for human judgment. That simple yet profound insight gave birth to Design Arena, a crowdsourced evaluation platform developed by the parent company Intelligence. Today, the platform boasts a staggering 5.3 million active global users who continuously rank, test, and critique AI-generated media.

On Monday, Intelligence announced a major milestone in its young history: a $7.9 million seed funding round led by venture capital heavyweight Index Ventures, with strategic participation from Conviction (led by Sarah Guo and Mike Vernal), A*, Valkyrie, and other prominent backers. Amid an industry increasingly plagued by compromised automated benchmarks and shifting consumer expectations, Intelligence has positioned itself as an indispensable gatekeeper for the next generation of AI models—turning human taste into a lucrative, enterprise-grade commodity generating $60 million in annualized recurring revenue (ARR).


Detailed Chronology: From Failed Game Engines to a Multi-Million Dollar Platform

The Genesis: The 2025 College Project

The story of Intelligence began not in a heavily funded Silicon Valley boardroom, but in a college dorm room just weeks before commencement in the spring of 2025. Grace Li and a tight-knit circle of engineering peers were attempting to build a novel AI game engine designed to democratize game development.

The technical hurdles of getting the underlying models to generate functional code and basic graphic assets were cleared relatively quickly. However, the resulting video games suffered from a glaring flaw: they lacked engagement. They were mechanically sound yet entirely unentertaining.

Recognizing that traditional algorithmic metrics were blind to the nuances of human entertainment value, Li and her co-founders shifted their focus away from game generation and toward the underlying problem itself: the complete absence of scalable human feedback loops specifically tailored for creative and generative design.

Pivoting to the Bottleneck

The team quickly realized that their dilemma was not isolated to game development. Across the entire generative AI landscape, frontier labs were spending billions of dollars training models to create websites, images, videos, and UI/UX designs. Yet, these labs had no reliable, high-volume mechanism to evaluate whether their outputs appealed to human aesthetics.

"It was the missing bottleneck for a lot of these models to make improvements in the design space," Li explains, reflecting on the pivot.

Rather than continuing to build the games themselves, the founders built an arena designed to evaluate creative media. Users were invited to test prompts, generating visual formats ranging from landing pages to complex graphic art. Through a streamlined "A vs. B" ranking interface, users evaluated outputs blindly, selecting which creation looked or performed better.

The pivot proved to be an immediate triumph. Just one week after opening the platform to the public, the startup closed its first major enterprise contract with a premier AI frontier lab. From that moment onward, the trajectory of Intelligence was cemented.

The $7.9M Seed Round and Current Scale

Scaling rapidly through viral adoption and organic discovery, Design Arena ballooned to an astonishing 5.3 million users worldwide. To handle the explosive infrastructural and computational demands of processing millions of human evaluations daily, the company sought institutional backing.

Monday’s announcement of a $7.9 million seed round led by Index Ventures provides Intelligence with the capital necessary to expand its enterprise offerings, reinforce its proprietary data pipelines, and scale its engineering teams. With prominent industry leaders like Sarah Guo and Mike Vernal joining the cap table through Conviction, the startup has solidified its status as one of the most closely watched infrastructure plays in the generative AI ecosystem.


Supporting Context & Metrics: Decoding the Business Model

How Design Arena Works

For the everyday consumer, navigating Design Arena feels remarkably similar to using an advanced model router. The interface features a clean, ChatGPT-style prompt window accompanied by intuitive dropdown menus allowing users to select their desired medium—ranging from full websites and digital illustrations to specialized UI mockups and graphic formats.

Once a user submits a prompt, format preference, and stylistic parameters, the platform bypasses single-model delivery. Instead, it prompts multiple underlying generative models to fulfill the request simultaneously. The user is then presented with a series of side-by-side "A vs. B" choices, iteratively ranking the competing outputs from best to worst.

The Enterprise Engine and Monetization

While the front-end user experience is straightforward and engaging, the true economic engine of Intelligence lies behind the enterprise curtain.

Participating AI labs plug into Design Arena to tap into an endless, real-time stream of human evaluation data. Because everyday users are largely indifferent to which specific model generated a given piece of media—their only objective is to receive the highest-quality output—their blind rankings produce unbiased, objective data regarding human preference.

This human-led evaluation data acts as the ultimate compass for model alignment. For frontier labs pouring massive resources into refining media-generating models, paying for this granular data is an absolute necessity. Consequently, Intelligence has unlocked a remarkably lucrative revenue stream, pulling in an impressive $60 million in ARR in less than a year of operation.

Geographic and Temporal Taste Tracking

Beyond simple ranking metrics, Intelligence leverages user authentication requirements to track deeper macroeconomic and cultural trends. Because users must log in to receive their generated outputs, the platform can meticulously analyze how visual tastes fluctuate over time and across different continents.

For instance, Grace Li notes distinct regional preferences within the platform’s analytics, observing that web dashboards originating from Asian markets frequently lean toward a more maximalist design philosophy compared to their Western counterparts.

These sophisticated qualitative datasets serve as a vital counterweight to traditional automated benchmarks. While automated testing frameworks operate at immense speeds and scales, they are inherently vulnerable to manipulation, gaming, and prompt leakage. This vulnerability was dramatically underscored last week during the high-profile security breach at Hugging Face, which laid bare the fragility of purely automated validation pipelines.


Official Statements and Industry Perspectives

The rapid rise of Intelligence and its human-in-the-loop evaluation model arrives at a fascinating crossroads for the artificial intelligence industry, where the limits of synthetic training data are becoming increasingly apparent.

Reflecting on the company’s explosive trajectory, Grace Li emphasizes the speed at which the market adopted their solution:

"It was the missing bottleneck for a lot of these models to make improvements in the design space. About a week later, we closed our first major deal with a frontier lab, and the rest is kind of history."

Industry analysts note that Intelligence’s success highlights a broader industry shift: as foundational text models reach a plateau of commoditization, the competitive frontier has decisively shifted toward multimodal generation, aesthetic alignment, and subjective quality control.

However, the path of crowdsourced human evaluation is not without its casualties. The sector remains intensely volatile, as evidenced by the high-profile collapse of Yupp AI, which shuttered its doors earlier this year despite raising an astronomical $33 million from prominent crypto and tech investors, including a16z crypto’s Chris Dixon. Despite boasting over 1.3 million users and securing several frontier models as enterprise clients, Yupp ultimately failed to construct a sustainable, long-term business model.

Conversely, other human-evaluation startups are commanding astronomical valuations. LM Arena, which applies a similar crowdsourced ranking methodology specifically to text-based conversational models, successfully secured a $150 million Series A valuation in January—achieving this milestone just four months after formally launching its paid enterprise product.

These divergent outcomes suggest that while human feedback is universally recognized as the holy grail of modern AI alignment, executing a sustainable, high-margin enterprise business model around it requires extreme precision, deep enterprise trust, and disciplined cost management.


Future Outlook: The Ongoing Quest for Machine Taste

As Intelligence deploys its fresh $7.9 million seed capital, the startup faces both immense opportunity and formidable challenges.

The demand for human preference data shows no signs of slowing down. As generative video, complex 3D rendering, and hyper-personalized UI generation become the next battlegrounds for tech giants and nimble startups alike, the need for reliable, un-gameable aesthetic evaluation will only intensify. Automated benchmarks will continue to serve as a baseline for computational efficiency, but the ultimate arbiter of quality will remain human consciousness.

By successfully bridging the gap between casual consumer interactions and enterprise-grade data pipelines, Intelligence has proven that human taste can be systematically captured, measured, and monetized. Whether the company can maintain its extraordinary $60 million ARR momentum and fend off looming competition in an increasingly crowded evaluation market will be one of the defining tech storylines to watch in the years ahead.

For Grace Li and her co-founders, what began as an undergraduate frustration over uninspired video games has successfully evolved into a foundational pillar of the global AI economy.


Editorial Disclosure: When you purchase products or services through links embedded in our articles, TechCrunch may earn a small commission. This commercial arrangement does not influence our editorial independence, investigative standards, or journalistic integrity.

Leave a Reply

Your email address will not be published. Required fields are marked *