Joint Hypothesis Problems in Life
In this post, I explore how a fundamental challenge in empirical asset pricing — the joint hypothesis problem — appears in many aspects of our daily lives. While we often take for granted that evaluation requires comparison, this simple principle leads to complications in how we assess everything from investment strategies to societal wellbeing. I thank Lucy Msall for comments.
Introduction
An important rule in life is that evaluation necessarily requires comparison. Whether we’re assessing investment performance, technological capabilities, or life experiences, we implicitly rely on some benchmark of what “good” looks like.
This fact is particularly relevant when individual evaluations must be aggregated into collective decisions. The solutions developed in finance, particularly the focus on “near-arbitrage” tests, might offer insights for this broader challenge of social decision-making.
Joint Hypothesis Problem in Asset Pricing
The challenge that begets evaluation was formally recognized in financial economics by Eugene Fama while developing the Efficient Market Hypothesis (EMH). Fama observed that when testing market efficiency, we’re actually testing two joint hypotheses:
- Markets efficiently incorporate available information
- Our model of expected returns is correct
In other words, if we find statistically significant excess returns, we can’t definitively say whether:
- The market is inefficient (hypothesis 1 fails)
- Our asset pricing model is misspecified (hypothesis 2 fails)
As a result, when we find that value stocks outperformed growth stocks, or small-cap stocks generated excess returns, we couldn’t definitively attribute these patterns to market inefficiency. These “anomalies” might instead reflect risk factors missing from our pricing models.
Joint Hypothesis Problem in Other Domains
This same fundamental challenge appears across various domains where evaluation depends on both the thing being evaluated and our model of what constitutes “good” performance.
Example 1: “LLMs are not that useful”
When someone claims Large Language Models aren’t useful, they’re simultaneously testing:
- The inherent capabilities of the LLM
- Their own ability to effectively prompt and utilize the tool
The evaluation combines both technological capability and user expertise. A seasoned prompt engineer might achieve remarkable results with the same LLM that appears useless to a novice user. This mirrors the asset pricing challenge — we can’t cleanly separate tool capability from usage expertise.
Example 2: “Prices went up so much”
Inflation experiences reflect two joint hypotheses:
- General price level changes in the economy
- Individual consumption patterns
Your personal inflation rate depends heavily on your consumption basket. A renter faces different price pressures than a homeowner. Someone who frequently dines out experiences different inflation than someone who mostly cooks at home. The "high inflation" narrative thus depends crucially on individual circumstances and choices.

Example 3: “I loved my PhD experience at Chicago”
Academic experiences combine:
- Objective program characteristics (curriculum, resources, faculty)
- Individual fit and expectations
The same program can be transformative for one student and challenging for another, depending on research interests, working style, and personal definition of a “good” academic environment. Like market efficiency tests, we can’t cleanly separate program quality from individual fit.
The Power of Near-Arbitrage Tests
In finance, researchers developed two approaches to address the joint hypothesis problem.
The first was to develop better benchmarks through multi-factor models. But the more innovative was the development of “near-arbitrage” tests – examining price patterns so extreme that they would indicate inefficiency under any reasonable model.
For instance, when identical assets trade at different prices in different markets, or when obvious public information isn’t reflected in prices, we don’t need a sophisticated asset pricing model to identify the inefficiency. These near-arbitrage tests became powerful tools precisely because they didn’t depend on specific assumptions about risk and return.
This insight offers a promising framework for addressing evaluation challenges in other domains. Instead of endlessly debating different benchmarks, we can look for evidence so compelling that it would indicate success or failure regardless of individual perspectives.
Example 1: “LLMs are not that useful”
Rather than debating LLM capabilities against varying user expectations, we might focus on:
- Examine whether LLMs can solve problems that are provably correct (like math)
- Look for internal consistency in responses rather than comparing to external benchmarks
- Clear violations of physical or computational constraints that would indicate deception
This mirrors financial near-arbitrage tests by identifying outcomes that would indicate capability (or lack thereof) under any reasonable evaluation framework. For example, an older version of ChatGPT incorrectly counted the number of “r” characters in the word “strawberry.”
Example 2: “Prices went up so much”
Economic policy already incorporates some near-arbitrage principles, though we could push further:
- We track “core necessities” inflation precisely because these items should be affordable under any reasonable social welfare function
- Minimum wage debates often center on whether basic survival is mathematically possible, a near-arbitrage test that transcends political views
- “Cliff effects” in welfare programs, where small income increases lead to net losses, represent welfare-reducing arbitrage opportunities we actively try to eliminate
Example 3: “I loved my PhD experience at Chicago”
Education assessment could benefit more from near-arbitrage thinking. While we do have some universal metrics:
- Faculty availability for extended periods
- Research resources necessary for the field
- Consistent performance of students to complete dissertations within reasonable timeframes
From Asset Pricing to Social Choice
The power of near-arbitrage tests in finance suggests a promising path forward for social decision-making. In fact, we have long identified certain principles that should hold regardless of individual preferences or circumstances. The U.S. Constitution’s inalienable rights represent early “near-arbitrage” tests: rights so fundamental that violating them would indicate social failure under any reasonable value system.
We see this principle at work in many established social institutions:
- Constitutional rights such as freedom of speech, due process, and equal protection
- Human rights frameworks such as access to clean water, protection from torture, child labor prohibitions
- Modern policy standards such as basic emergency medical care or access to education
This near-arbitrage framework becomes particularly valuable for emerging challenges like inequality, climate change, or technological regulation. Just as financial economists look for price patterns that would indicate inefficiency under any reasonable model, we can identify social outcomes that would indicate policy failure regardless of individual circumstances or preferences.
Of course, this framework doesn’t solve all our evaluation challenges. But by focusing on near-arbitrage tests – evidence so compelling it transcends individual benchmarks – we might find more constructive ways to make collective decisions in a diverse society. Just as financial economics moved beyond endless debates about the “right” asset pricing model (or did we?), social choice theory might benefit from seeking evidence that holds across all reasonable social welfare functions.
Conclusion
In an increasingly polarized world, the ability to find common ground becomes ever more crucial. The near-arbitrage framework from finance offers a powerful tool: instead of trying to resolve fundamentally different perspectives, we can look for evidence so compelling that it transcends individual benchmarks.
The joint hypothesis problem thus offers not just a challenge, but an opportunity. By pushing us to seek near-arbitrage tests, it might help us find common ground even in our most contentious social debates. Just as financial markets become more efficient when we identify clear arbitrage opportunities, our social decisions might improve when we focus on evidence that holds regardless of individual circumstances.