Game Theory Meets Machine Learning:
Designing Auction-Based Markets for Online
Advertising
References:
Athey and Ellison, QJE 2011
Athey and Nekipelov, working paper, 2011 and WIP
Athey’s Fisher Shultz Lecture, 2011
Roadmap
Applying Economic
Methodology to
Online Ads
A “classic” economic
approach in detail
Current directions
Economic v. ML
Methodology
What are economists
trying to do and why?
Can ML take anything
from econometrics?
Can econometrics take
anything from ML?
Structural Economic Modeling v.
“Classic” Machine Learning
• Economist focuses on “counterfactual predictions”
• Uncover “primitive” preferences and use behavioral assumptions, e.g.:
• Estimate demand for differentiated products using consumer choice data • Assume firms maximize profits
• Infer costs that rationalize prices chosen
• Predict how profit-maximizing prices change if firms merge
• Joint dist’n of features and outcomes changes with policy
• Make predictions outside the setting you have observed
• Emphasis on causality in structural modelling
• Two models might be equally good in current environment
• A structural model where estimates have a causal interpretation does better in very different environments (not just overfitting issue)
• Techniques for using only the “good” variation in input features, and
ignoring “bad” variation
Alternative Empirical Economics Methods and
Their Applications in Market Design
Field experiments
Compare design alternatives
Estimation of models
Create exogenous variation
Generate data from
alternative/new designs
Model validation
Observational Data and
“Reduced Form” Models
Use “natural” or “quasi”
experiments, emphasize solving
causality problem
Statistical techniques to only use
“good” variation: e.g. IV
Estimate effect of policy that
occurred; limited counterfactuals
Structural models
Make behavioral assumptions
Profit maximization
“Rational” decision-making
Observational or
quasi-experimental data used to
estimate economic primitives
Used for evaluating market design
and algorithms
Impact of policies on welfare/profit to predict long-run responses
Efficiency v. revenue tradeoffs Consider market designs not currently in use
Interdisciplinary Learnings
• Econometrics• Bayesian v. Classical
• Less emphasis on standard errors, more on model specification and selection
• More emphasis on testing/validating structural models using experiments or multiple regime changes
• Offline: market design: Athey/Haile (2002), Athey/Levin/Seira (2011), Athey/Coey/Levin (forthcoming)
• More emphasis on learning objectives rather than assuming behavior
• Athey/Nekipelov (in progress, today’s talk)
• Data driven model selection and
segmentation within economic frameworks • Large number of features
• Establish formal properties of some ML methods (but clearly define problem first) • See Hal Varian’s recent work for time
series/prediction issues…
• Machine learning
• How to make the best use of observational data to get causal estimates that give accurate predictions in new environments • Instrumental variables in machine
learning models
• How to incorporate in scalable models
• Not enough variation to be fully identified
• How to ensure you produce usable and reasonable estimates for every unit
Experiments and Models in Internet Search:
Click Prediction Case Study
Algorithms: structural models that predict clicks for alternative rankings of content Database of potential content Metrics about user interaction with content Algorithm A Exploration
Real-time model update
Scores Rankings of content User clicks Algorithm B User clicks versus Advertiser receives clicks and pays PPC based on scores
Advertiser updates bids
Scores Rankings of content SHORT TERM EXPERIMENT LONG TERM RESPONSE Use Structural Model
Simple Example: User Model of Clicks
Problem with observational dataMore “clickable” ads in high positions Ad ranking respond to indiv. users
Experiment only: expensive, not sufficient
Too many permutations of ads
Dedicated experiment + model
For new, very different algorithm
Practice (industry lore)
Some exploration
Pool observational and experimental data for estimation
Proposal: re-use existing experiments + model
Any experiment that generates exogenous variation in ad rankings is a candidate
User Model of Clicks:
Results from Spring 2009 Experiments
Clicks as a Fraction of Top First Position Clicks Search
phrase: iphone viagra
Model: OLS IV OLS IV
Top Position 2 0.66 0.67 0.28 0.66 Top Position 3 0.40 0.55 0.14 0.15 Side Position 1 0.04 0.39 0.04 0.13
OLS includes advertiser effects and position effects
IV estimates show smaller position impact than OLS, as expected. We
interpret this as the causal effect of position.
Since position discounts affect advertiser bid shading, this is very important
for counterfactual advertiser models.
Auction-based Platforms
Platform markets
Media markets, credit cards, dating, video games, operating
systems…
Two groups of customers, externalities (typically indirect
network effects)
Platform can help internalize externalities
Auction-based platforms
Online advertising
Used cars
Market Design Insights
“Size of the pie” v.
distribution of rents
Short-run distribution affects
long run size of pie and
platform revenue
Direct: platform share of pie
Indirect: participation
Indirect channel often
dominates
Policy conclusion: design
markets for participation
Preference policies for weak
bidders may not hurt
revenue
Internet search advertising
Substantial tradeoffs between
efficiency and short-run
revenue
Incentives to price discriminate and favor market thickness over efficiency
Incentives to show too many irrelevant ads
Incorporating long-run bidder
and user responses leads to
different design choices
Need for models
Understand objectives and
welfare of participants
A structural model of search advertising
• This is just classic monopsonist problem
• Can estimate these quantities from search engine data
by simulating impact of hypothetical bid changes
Estimates of AC(q), MC(q), and implied value for
a high-value search phrase
Estimates of AC(q), MC(q), and implied value for
a high-value search phrase
Counterfactual Analysis
• Once values known, can recompute equilibrium
• Use assumption on advertiser objectives
• See paper for existence/uniqueness, computational algorithm
• Homotopy method
• Making eqm computation scale up effectively is an open problem
Example: Score coarsening
Scores correspond to search engine’s estimate of consumer’s
propensity to click
User-level data is important
What if we lose ability to see some user characteristics?
Should we use our most accurate and personalized click prediction
algorithm in the auction, or a coarsened version?
Model this by representing log-score as sum of two components
and remove one of them
Economic tradeoff:
Personalized click predictor implies search engine estimates low click
propensity for less preferred ads, high click propensity for more preferred
ads
In click-weighted auction, user’s less preferred ads are no longer effective
competitors in the auction
Reducing Click Prediction Accuracy Has Competing Effects: Reduced
Welfare, Increased Revenue
Crucial role for advertiser modeling:
Short and long run revenue predictions go in opposite directions.
Algorithm:
Before
Coarsening
After
Coarsening
After
Coarsening
Bids:
Original
Original
Counterfactual
Welfare
6.956
6.903
6.715
Revenue
2.091
2.036
2.138
Advertiser
Profit
4.865
4.867
4.765
Revenue falls in short run experiment Revenue rises when bids adjustIssues with Applying Model in Practice:
Advertiser Heterogeneity
Test model performance when system changes
Estimate values before change
Predict bids after change
Heterogeneity in attention/engagement
Many advertisers change their bids rarely, don’t monitor closely
Heterogeneity of objectives
Value from clicks, impressions, position
Budget constraints
Heterogeneity of tools to implement objectives
Agencies
Tools providers
Manual optimization
Initial Evidence of Heterogeneity
Sample of campaigns, not weighted by importance
Estimate separately for
Bayesian Approach to Heterogeneity
Identification (Causality)
Econometric approach
Write down formal assumptions and prove theorem that says that with large enough dataset, causal parameters can be “identified”
Identification: Objective type and parameters
Observe each advertiser over a time period with “system shocks” Change in pricing or ranking algorithms
The algorithm will change outcomes differently for each search query
Shocks are heterogeneous across advertisers, and also heterogeneous across different keywords of a given advertiser
Assume timing of system shocks uncorrelated with preference shocks Initial implementation: preference parameters constant over time
Complications: system changes prior to holidays
Different objectives imply different response to shock
Time and attention process
Observe frequency of bid change together with estimated objectives
Relationship between bid change and potential returns to bid change based on each objective
Inputs of Machine Learning
Tree-based models: map behavioral types onto segments
E.g. “Finance bidders are typically profit maximizers”
In an environment with lots of advertiser observables
Many advertiser “exogenous” characteristics
Multi-dimensional behavioral description (posteriors on types, time cost/responsiveness)
What are the most relevant segmentations for predicting advertiser
responsiveness and objective types?
“Training” and “Testing”
Test performance of model for system changes outside training set
Look for system changes qualitatively different than training changes
Conclusions
Econometric methods have a lot to add to market design
and marketplace management
Field experiments and structural models combine
Bayesian and ML methods rather than “traditional”
classical econometrics
Structural Models
Field Experiments
Offline
Used in research, occasional policy evaluations for market design, e.g.:• Optimal reserve price modeling for softwood lumber trade dispute (Athey & Ingraham)
• Comparing auction format in treasury auctions (e.g. Hortacsu, Kastl), timber (Athey, Levin, Seira; Athey, Coey, Levin)
• Detection and damages in bidder collusion (Marshall et al, Bajari, Porter, Pesendorfer, etc.)
• Mergers in auction markets (Froeb)
Expensive, slow
Relatively rare but can be very influential
Online
Algorithms used are structural models Long run impact of market design, algorithms: where experiments are expensive, slow, uninformative• E.g. advertiser reactions (Ostrovsky & Schwarz, Athey & Nekipelov)
Essential and integral to businesses • Algorithm training and real-time
explore/exploit
• Selection and evaluation of new algorithms: REQUIRED
• Tuning algorithms and UI
• Validation of structural models Researchers can create (e.g. eBay) or exploit (Einav & Levin)
GSP v. Vickrey: Revenue and Efficiency
Crucial role for advertiser modeling:
Need to recover advertiser valuations to predict outcomes under Vickrey.
Search Phrase 1
Search Phrase 2
GSP
Vickrey
GSP
Vickrey
Welfare
8.228
8.231
6.513
6.531
Revenue
2.490
1.947
3.114
3.153
Vickrey auction: bidders pay “expected externality,” bid truthfully • Allocation is efficient in every instance of a user query
GSP has asymmetric bid shading
• Bidders with close competitors bid close to value, others shade a lot
• For some score shock realizations, low-value bidder beats high value bidder • Score shock variation also breaks revenue equivalence result
Estimation
Markov Chain Monte Carlo
Output of model
Posterior distributions over objective types and parameters
Counterfactual estimates
Issue: equilibrium with heterogeneous objectives
Bidders have same posterior as econometrician over opponent types
◻ But why do bidders know the same? Why not less, or more?
Other approaches??
Posterior distributions over advertiser profit, revenue, total
welfare
Estimation Algorithm:
Computing Q(b) and TE(b) Curves
Historical
Auctions Simulation with Counterfactual
shocks to scores and entry
New pretend bid
Q(bid)
TE(bid)
Actual bid Actual bidQ(bid)
TE(bid)
Estimation Details
Athey-Nekipelov (2011)
Estimate distributions of score shocks nonparametrically
Brute force simulation, numerical derivatives of Q() and TE()
Estimates have desirable properties (uniform convergence,
asymptotical normality), asymptotic variance formulas verified using
Monte Carlo
Athey-Nekipelov (in progress)
Assume score shocks are jointly log-normal (good approx)
Derive analytic formulas for derivatives in terms of integrals of
normal r.v.’s
Use numerical approximations for derivatives
Greatly improved robustness and accuracy
Both methods implemented “at scale” on millions of
advertiser bids in parallel
Require replication and aggregations, no numerical optimization or
convergence
More Accurate Click Prediction Increases
Efficiency, But Can Decrease Revenue
Per-Click Bid Estimated Revenue Bid (normalized to 1st position)
Price Per Click Platform Revenue (including position discounts) Expected Clicks (including position discounts)
“Coarse” click predictor with two bidders with equal bids, avg. clickability
b b s b b s s
b b s R/s α2 R α2 s
“Granular” click predictor identifies user types:
Half of users like A better so true score is (1+d) s for A and (1−d) s for B Half of users like B better so true score is (1+d) s for B and (1−d) s for A
b b (1+d) s b(1−d) /(1+d) b(1−d) s (1+d) s
b b (1−d) s R/((1−d) s) α
2 R α2(1−d) s
Differences in outcomes: “Granular” – “Coarse”