Media & Press

Huxiu Think Tank Services | The 'Involution' of Large Models Ends, the Second Half of Enterprise AI in the Subjective World Model

At WAIC, Tezign proposed the Subjective World Model (SWM), a four-layer architecture modeling the real decision logic of consumers, linking GEA enterprise intelligent agents to business operations, with exclusive in-depth interviews building high barriers, opening a new track for enterprise AI.

Category

Media & Press

Date

2026-07-21

Read Time

8 min read

The 2026 World Artificial Intelligence Conference (WAIC) opened on July 17 in Shanghai. This year's conference is the largest ever, with Agents becoming the most concentrated narrative axis on the industrial side—from computing power infrastructure to autonomous execution systems, discussions on 'what AI can do' have been quite thorough.

Beyond the mainstream narrative of Agents, we discovered an interesting angle that may represent another direction for AI development—the Subjective World Model. This concept also appeared in the CCTV news studio, proposed by Professor Fan Ling from Tongji University, who also founded Tezign.

The characteristic of this model is that it allows AI to begin to understand people—not just generating language, but modeling the psychological structure, decision logic, and behavioral tendencies of a specific individual. He refers to the system that Tezign is researching as the 'Subjective World Model' (SWM).

This concept subsequently sparked intense questioning at the WAIC booth, showing significant differences from the familiar LLM paradigm.

The reporter visited the booth to explore further, and this article aims to systematically outline the technical path, data architecture, and competitive barriers of SWM.

Structural Boundaries of LLM

In the past three years, the penetration speed of large language models in enterprise applications has exceeded expectations. Content generation, code assistance, knowledge retrieval, customer service automation—these scenarios share a common feature: the core task is language output, rather than understanding a specific person.

When the application scenario shifts from 'generating reasonable text' to 'understanding consumer decisions', the structural boundaries of LLM begin to emerge.

The pre-training objective of LLM is next-token prediction: predicting the probability distribution of the next token given the context. After training on trillions of tokens, the model has established a high-quality approximation of the statistical laws of human language. This enables it to fluently complete language-level tasks, but it models language itself, rather than the psychological structure of the speaker behind the language.

This boundary is particularly prominent in brand decision scenarios: consumers' purchase motivations, value weights, risk preferences, and the systematic bias between self-reported and actual behavior are all outside the modeling objectives of LLM. User profiles generated by LLM are essentially projections of language statistics, rather than modeling consumer psychology.

Technical Path of the Subjective World Model

The Subjective World Model (SWM) is an independently designed model architecture by Tezign Technology to address the aforementioned boundaries, serving as the underlying technical framework for its consumer research platform Atypica.

The core proposition of SWM is to treat 'modeling a person's subjective world' as an independent training objective, rather than fine-tuning or domain adaptation of LLM. Its training objectives, data systems, and evaluation methods are all independent of the LLM system.

Architecture Design: Four-Layer Collaborative Modeling

SWM breaks down the 'subjective world of consumers' into four independently modelable and collaboratively inferable layers:

Expression Layer

The training data consists of billions of original social media posts, with the modeling objective being the high-dimensional mapping relationship between language style and demographic/psychographic signals.

The same consumer perspective presents systematically different language encodings across different groups—word frequency distribution, emotional intensity, and implicit value judgments all shift predictably with identity variables. The task of the expression layer is to infer the identity structure and psychological characteristics of the speaker from the surface of language, providing a basis for individual identification for subsequent layers.

Story Layer

The training data consists of tens of thousands of hours of one-on-one in-depth interviews, each lasting 1-2 hours, generating 5,000-20,000 words of unstructured data. The modeling objective is the temporal causal structure of behavioral motivations.

There is a fundamental difference in information density between questionnaire data and in-depth interview data: questionnaires capture static preference vectors, while in-depth interviews capture motivation chains—the complete temporal structure of how individuals encounter products, form consideration intentions, trigger decisions, and the self-explanation mechanisms after purchase. The latter contains causal information that the former cannot extract.

The data from the story layer has threefold non-replicability: first, it cannot be crawled—there is a lack of such high-density, one-on-one consumer behavior narratives on the internet; second, it cannot be synthesized—LLM-generated simulated interview data will systematically distort in cross-validation at the cognitive and behavioral levels; third, it cannot be rushed—data accumulation relies on long-term customer service relationships and professional research capabilities, with Tezign's story layer data coming from consumer research projects of over 180 enterprise clients over ten years.

Cognition Layer

The training data includes behavioral judgment questionnaires and standardized psychological scales (Schwartz Value Scale, BFI Personality Scale, etc.), with the modeling objective being the individual's true value weight system and risk preference coefficients.

The core basis for the cognition layer is self-report bias—there is a measurable systematic gap between individuals' expressed preferences and their true decision weights. The cognition layer is used to model and correct this gap, allowing the model to not rely on consumers' self-reports but to restore their true decision weight structure.

Behavior Layer

The training data consists of economic game experiments (ultimatum game, public goods game, etc.) and real transaction records, with the modeling objective being behavioral economics parameters: loss aversion coefficient, temporal discounting rate, social norm sensitivity, etc.

The behavior layer addresses the problem of cross-situation transfer: when the model encounters new scenarios not present in the training data, it needs to infer responses based on stable parameters at the individual level across situations, rather than degrading to a general description.

Why Synthetic Data Replacement is Not Feasible

'Using LLM to synthesize in-depth interview data to replace real data collection' is a direct idea to reduce the cost of story layer data and is the core direction of active defense in SWM architecture design.

Synthetic data performs well in validation at the expression layer—LLM-generated text has a high degree of simulation capability in language style. However, in cross-validation at the cognition and behavior layers, synthetic data will systematically fail.

The narratives of real respondents have cross-situation consistency: the core value weights of individuals (such as risk aversion tendencies) will be expressed in different forms across different purchasing scenarios, forming an internally consistent 'psychological signature'. LLM does not maintain a continuous internal state during the generation process, but independently predicts the most reasonable token at each generation—therefore, synthetic 'respondents' lack consistency in value weights across scenarios.

The cross-validation mechanism of the cognition layer will detect and filter out this inconsistency, preventing synthetic data from entering model training. This mechanism also ensures that the competitive barrier of real data accumulation cannot be bypassed by synthetic means.

Quantitative Indicators and Actual Deployment

After four-layer collaborative training, SWM outputs AI Persona—a digital individual that can maintain an internally consistent psychological structure, value weights, emotional responses, and decision tendencies under any new prompt.

Current publicly available quantitative results include:

- Behavioral simulation accuracy: 85% (compared to real in-depth interview benchmarks)

- Coverage scale: 300,000+ AI Personas (social data sources), 10,000+ high-precision AI Personas (in-depth interview data sources)

- Delivery cycle: < 30 minutes (compared to traditional consumer research of 4-8 weeks)

SWM is currently the underlying architecture of Atypica (Tezign Technology's insight research intelligent agent product) and has completed deployment verification in the real business scenarios of over 180 enterprise clients globally, covering industries such as fast-moving consumer goods, technology, automotive, and retail.

Collaboration Between SWM and GEA Architecture

SWM addresses the understanding side issue—structural modeling of the consumer's subjective world. The realization of the commercial value of consumer insights relies on its effective transformation into business execution. Who is responsible for this transformation?

Tezign GEA is the answer that specifically undertakes this transformation link.

GEA is Tezign's self-developed four-layer enterprise-level intelligent agent architecture: the intention layer is responsible for structured parsing of business objectives; the orchestration layer is driven by Tezign's self-developed divergent reasoning model (Creative Reasoning Model), capable of coordinating 30+ foundational models in a single task; the skills layer provides 400+ modular professional skills; the context layer (Context System) accumulates enterprise knowledge, brand DNA, and historical insights, continuously callable.

The consumer insights output by SWM, combined with the Context System and the enterprise's historical data, form an end-to-end closed loop from consumer understanding to business execution. GEA has been deployed in over 50 countries and regions worldwide, with a monthly token deployment volume exceeding 10 billion.

Structural Analysis of Competitive Barriers

From the perspective of competitive landscape, the barriers of SWM are composed of four layers:

Data barrier: The ten-year accumulation of in-depth interview data in the story layer is irreplaceable, non-crawlable, non-synthesizable, and non-rushable, forming the most core single-point barrier.

Validation barrier: The cross-validation mechanism of the cognition and behavior layers systematically blocks the alternative paths of synthetic data, ensuring the sustainability of the data barrier.

Scale barrier: The coverage scale of 300,000+ AI Personas allows a single study to achieve a level of granularity that traditional methods cannot achieve under time and cost constraints.

Deployment barrier: The verification accumulation of over 180 enterprise clients in real business scenarios has formed a first-mover advantage in industry coverage depth and scenario density.

Conclusion

The technological generational shifts in consumer research—from focus groups to questionnaire platforms, from NLP to sentiment analysis—have all been advancements in the dimension of 'more efficiently collecting what consumers said'. SWM attempts to transcend this dimension, directly modeling 'why consumers do this, and what they will do next'.

From a technical path perspective, this is currently feasible. From the data barrier perspective, the advantages accumulated over the past decade have already formed. Atypica has also validated the feasibility of the complete productization of this technical framework.

The commercialization process of AI supply-side capabilities is still accelerating, and the next differentiation dimension of enterprise AI competition will depend on who understands consumers better—this understanding capability cannot be API-ized, cannot be platformed, but can only be accumulated.

Readers interested in this direction can visit the Tezign booth (H1-C135) during WAIC to experience the demo of Atypica on-site; direct interaction with the product is the most straightforward way to understand the actual effects of SWM. The exhibition runs until July 20.

(Source: WeChat Official Account Huxiu Think Tank Services)

Related Recommendations

Tezign Technology Selected for 2026 'China AGI Innovation Impact Institutions TOP 30'
Media & Press2026-08-05

Tezign Technology Selected for 2026 'China AGI Innovation Impact Institutions TOP 30'

Tezign Technology's Enterprise-Level Intelligent Agent Sets a Benchmark, Winning the 2026 Extraordinary Award for Best Product and Case Award
Media & Press2026-07-23

Tezign Technology's Enterprise-Level Intelligent Agent Sets a Benchmark, Winning the 2026 Extraordinary Award for Best Product and Case Award

QbitAI | Watch the WAIC Model 'Mind Reading'! It's on fire on site!
Media & Press2026-07-22

QbitAI | Watch the WAIC Model 'Mind Reading'! It's on fire on site!

Ready when you are

Put enterprise agents to workon a real business problem.