消费者最优与对偶三角Consumer optimum and the duality triangle
Utility maximisation, duality, and the Slutsky decomposition
整个需求理论只有一个优化问题,从四个角度看就成了四组结论。
Demand theory contains exactly one optimisation problem. Viewed from four angles it becomes four sets of results.
核心方程Core equations
展开逐步推导(3 步)Show the 3-step derivation
- 写拉格朗日并取一阶条件Set up the Lagrangian and take first-order conditions$$\mathcal{L}=U(x,y)+\lambda(I-p_xx-p_yy);\quad U_x=\lambda p_x,\ \ U_y=\lambda p_y$$两式相除消掉 \(\lambda\),得 \(MRS=\dfrac{U_x}{U_y}=\dfrac{p_x}{p_y}\)——主观替代率等于市场替代率。Dividing one condition by the other cancels \(\lambda\) and gives \(MRS=\dfrac{U_x}{U_y}=\dfrac{p_x}{p_y}\): the subjective rate of substitution equals the market rate.
- 解出马歇尔需求Solve for Marshallian demand$$x=x(p_x,p_y,I),\qquad y=y(p_x,p_y,I)$$代回效用得间接效用函数 \(V(p_x,p_y,I)\):给定价格与收入能达到的最高效用。Substituting back into utility yields the indirect utility function \(V(p_x,p_y,I)\): the highest utility attainable at given prices and income.
- 对偶问题给出补偿需求与支出函数The dual delivers compensated demand and the expenditure function$$x^c(p_x,p_y,\bar U),\qquad E(p_x,p_y,\bar U)$$\(V\) 与 \(E\) 互为反函数:\(E(p,V(p,I))=I\)、\(V(p,E(p,\bar U))=\bar U\)。这是对偶的全部内容。\(V\) and \(E\) are inverses of one another: \(E(p,V(p,I))=I\) and \(V(p,E(p,\bar U))=\bar U\). That is all duality amounts to.
展开逐步推导(2 步)Show the 2-step derivation
- Shephard 引理来自包络定理Shephard's lemma follows from the envelope theorem$$\frac{\partial E}{\partial p_x}=\frac{\partial}{\partial p_x}\bigl[p_xx^c+p_yy^c\bigr]=x^c$$只对参数求偏导,内生量的变化不用管——这正是包络定理。多算链式项是这里的头号错误。Differentiate with respect to the parameter only; the response of the endogenous variables can be ignored. That is precisely what the envelope theorem buys you, and adding the chain-rule terms anyway is the commonest error here.
- Roy 恒等式来自对 V 全微分Roy's identity follows from totally differentiating V$$\frac{\partial V}{\partial p_x}=-\lambda x,\qquad \frac{\partial V}{\partial I}=\lambda$$两式相除,\(\lambda\) 约掉。负号来自价格上升使效用下降,写掉负号会得到负的需求。Divide one by the other and \(\lambda\) cancels. The minus sign comes from the fact that a price rise lowers utility; drop it and demand comes out negative.
展开逐步推导(3 步)Show the 3-step derivation
- 从恒等式 xᶜ(p, Ū) = x(p, E(p, Ū)) 出发,对 pₓ 求导Start from the identity xᶜ(p, Ū) = x(p, E(p, Ū)) and differentiate with respect to pₓ$$\frac{\partial x^c}{\partial p_x}=\frac{\partial x}{\partial p_x}+\frac{\partial x}{\partial I}\cdot\frac{\partial E}{\partial p_x}$$右边第二项用了链式法则:\(E\) 也依赖 \(p_x\)。The second term on the right uses the chain rule, since \(E\) depends on \(p_x\) as well.
- 把 Shephard 引理代进去Substitute Shephard's lemma$$\frac{\partial E}{\partial p_x}=x^c=x \quad(\text{equal at the initial point})$$补偿需求与马歇尔需求只在初始点重合,这一步是整个推导的枢纽。Compensated and Marshallian demand coincide only at the initial point. This step is the pivot of the whole derivation.
- 移项Rearrange$$\frac{\partial x}{\partial p_x}=\frac{\partial x^c}{\partial p_x}-x\frac{\partial x}{\partial I}$$替代效应恒非正(凹性保证);收入效应符号取决于正常品还是低档品。吉芬品要求低档且收入效应压过替代效应。The substitution effect is never positive (concavity guarantees it); the sign of the income effect depends on whether the good is normal or inferior. A Giffen good requires an inferior good whose income effect outweighs the substitution effect.
交互图Interactive chart
教学算例Worked example
\(u=\alpha\ln q+(1-\alpha)\ln Q,\ \alpha=0.3,\ y=100,\ p_y=5\);咖啡价格 \(p_x\) 由 2 涨到 4。
\(u=\alpha\ln q+(1-\alpha)\ln Q,\ \alpha=0.3,\ y=100,\ p_y=5\); the price of coffee \(p_x\) rises from 2 to 4.
| 原需求 \(q_0=\alpha y/p_x\)Initial demand \(q_0=\alpha y/p_x\) | 15.000 | 份额固定:花在咖啡上的钱恒为 30The share is fixed: spending on coffee stays at 30 |
| 新需求 \(q_1\)New demand \(q_1\) | 7.500 | 价格翻倍、支出不变 ⇒ 数量减半Price doubles, spending unchanged, so quantity halves |
| 原效用 \(v_0\)Initial utility \(v_0\) | 2.659755 | \(=\ln y-\alpha\ln p_x-(1-\alpha)\ln p_y+\alpha\ln\alpha+(1-\alpha)\ln(1-\alpha)\)\(=\ln y-\alpha\ln p_x-(1-\alpha)\ln p_y+\alpha\ln\alpha+(1-\alpha)\ln(1-\alpha)\) |
| 补偿需求 \(q^c\) @ 新价Compensated demand \(q^c\) at the new price | 9.233583 | 把收入补到能维持 \(v_0\) 时会买的量What she would buy if income were topped up to hold \(v_0\) |
| 替代效应Substitution effect | −5.766 | 沿同一条无差异曲线滑动,恒为负A slide along one indifference curve, always negative |
| 收入效应Income effect | −1.734 | 实际购买力下降;咖啡是正常品故同向Real purchasing power falls; coffee is normal, so the effect runs the same way |
| 总效应Total effect | −7.500 | −5.766 + (−1.734),斯勒茨基恒等式成立−5.766 + (−1.734); the Slutsky identity holds |
经典文献Original sources
- Slutsky, E. (1915). "Sulla teoria del bilancio del consumatore." Giornale degli Economisti.替代–收入分解的原始文献,被埋没近 20 年才由 Hicks 与 Allen 重新发现。The original source of the substitution–income decomposition. It lay buried for nearly twenty years until Hicks and Allen rediscovered it.
- Hicks, J. R. & Allen, R. G. D. (1934). "A Reconsideration of the Theory of Value." Economica.把序数效用与替代弹性带进主流,Slutsky 分解由此广为人知。Brought ordinal utility and the elasticity of substitution into the mainstream, and with them the Slutsky decomposition.
- Roy, R. (1947). "La distribution du revenu entre les divers biens." Econometrica 15, 205–225.Roy 恒等式。Roy's identity.
- Shephard, R. W. (1953). Cost and Production Functions. Princeton University Press.Shephard 引理,原本是为成本函数写的——见模型 2,两者是同一件事。Shephard's lemma, originally written for cost functions — see Model 2, where it is the same result.
易错点Where it goes wrong
- 把补偿需求与马歇尔需求全程当成相等。它们只在初始点重合。斯勒茨基推导的枢纽正是「在初始点 \(x^c=x\)」这一步,用早了或用晚了结论都会错。Treating compensated and Marshallian demand as equal throughout. They coincide only at the initial point. The pivot of the Slutsky derivation is exactly the step where \(x^c=x\) holds; invoke it too early or too late and the result is wrong.
- Roy 恒等式忘负号。\(\partial V/\partial p_x\lt0\),分母 \(\partial V/\partial I\gt0\),不加负号会得到负需求。Dropping the minus sign in Roy's identity. \(\partial V/\partial p_x\lt0\) while \(\partial V/\partial I\gt0\); without the minus sign demand comes out negative.
- 用包络定理时多算链式项。\(\partial E/\partial p_x\) 只对显式出现的 \(p_x\) 求导,\(x^c\) 随 \(p_x\) 的变化不用管——这是包络定理的全部好处,多算就白用了。Adding chain-rule terms when using the envelope theorem. \(\partial E/\partial p_x\) differentiates only the \(p_x\) that appears explicitly; the response of \(x^c\) to \(p_x\) can be ignored. That is the entire benefit of the theorem, and adding the terms throws it away.
生产、成本与 CES 替代弹性Production, cost duality and the CES family
Production, cost duality, and the CES family
一个参数 σ 就把 Leontief、柯布–道格拉斯、线性生产函数串成一条谱。
One parameter, σ, strings Leontief, Cobb–Douglas and linear technology onto a single spectrum.
核心方程Core equations
展开逐步推导(3 步)Show the 3-step derivation
- 成本最小化的一阶条件First-order condition for cost minimisation$$\frac{q_K}{q_L}=\frac{r}{w}\ \Longrightarrow\ \frac{a}{b}\Bigl(\frac{K}{L}\Bigr)^{\rho-1}=\frac{r}{w}$$\(RTS\) 等于要素价格比——与消费者的 \(MRS=p_x/p_y\) 一模一样。The \(RTS\) equals the factor price ratio, exactly as \(MRS=p_x/p_y\) does for the consumer.
- 两边取对数,把指数拿下来Take logs of both sides to bring the exponent down$$(\rho-1)\ln\frac{K}{L}=\ln\frac{r}{w}-\ln\frac{a}{b}$$用 \(\ln x^n=n\ln x\)。这一步是为了让下一步能直接读出弹性。Using \(\ln x^n=n\ln x\). The point of this step is to let the next one read the elasticity straight off.
- 解出对数比值,再对 ln(w/r) 求导Solve for the log ratio, then differentiate with respect to ln(w/r)$$\ln\frac{K}{L}=\frac{1}{\rho-1}\Bigl(\ln\frac{r}{w}-\ln\frac{a}{b}\Bigr) \ \Longrightarrow\ \frac{d\ln(K/L)}{d\ln(w/r)}=\frac{1}{1-\rho}=\sigma$$\(\ln(r/w)=-\ln(w/r)\),负号与 \(\rho-1\) 的负号相消。σ 的定义就是这个对数导数,不是别的。\(\ln(r/w)=-\ln(w/r)\), and that minus sign cancels the one in \(\rho-1\). σ is defined as this log derivative and as nothing else.
展开逐步推导(2 步)Show the 2-step derivation
- ρ→0 的极限是柯布–道格拉斯The ρ→0 limit is Cobb–Douglas$$\lim_{\rho\to0}\bigl(aK^{\rho}+bL^{\rho}\bigr)^{1/\rho}=K^{a}L^{b}\quad(a+b=1)$$先确认是 \(1^{\infty}\) 型未定式,取对数化成 \(0/0\) 再用洛必达。跳过这一步直接写结论是这节最常见的偷工。Confirm first that this is a \(1^{\infty}\) indeterminate form, take logs to turn it into \(0/0\), then apply LHôpital. Skipping straight to the answer is the standard shortcut taken here — and the standard place marks are lost.
- ρ→−∞ 给 Leontief,ρ→1 给线性ρ→−∞ gives Leontief, ρ→1 gives the linear case$$\sigma=\tfrac{1}{1-\rho}:\quad \rho\to-\infty\Rightarrow\sigma\to0;\quad \rho\to1\Rightarrow\sigma\to\infty$$σ 的经济含义:要素价格比变动 1%,要素投入比变动 σ%。σ=0 表示完全不可替代。What σ means economically: a 1% change in the factor price ratio changes the factor input ratio by σ%. σ=0 means the factors cannot be substituted at all.
交互图Interactive chart
教学算例Worked example
\(a=b=\tfrac12\)。工资租金比 \(w/r\) 翻倍,问资本–劳动比 \(K/L\) 变多少。
\(a=b=\tfrac12\). The wage–rental ratio \(w/r\) doubles: by how much does \(K/L\) change?
| ρ = −1 ⇒ σ = 0.5ρ = −1 ⇒ σ = 0.5 | ×1.414 | \(2^{0.5}\):替代困难,比例只涨 41%\(2^{0.5}\): substitution is hard, so the ratio rises only 41% |
| ρ = 0 ⇒ σ = 1ρ = 0 ⇒ σ = 1 | ×2.000 | 柯布–道格拉斯:等比例替代Cobb–Douglas: proportional substitution |
| ρ = 0.5 ⇒ σ = 2ρ = 0.5 ⇒ σ = 2 | ×4.000 | \(2^{2}\):替代容易,比例翻两番\(2^{2}\): substitution is easy, so the ratio quadruples |
| ρ → 0 的极限The ρ → 0 limit | \(\sqrt{KL}\) | sympy 验证:CES 确实退化成柯布–道格拉斯Verified in sympy: CES does collapse to Cobb–Douglas |
经典文献Original sources
- Cobb, C. W. & Douglas, P. H. (1928). "A Theory of Production." American Economic Review 18(1), 139–165.柯布–道格拉斯函数的原始文献,用 1899–1922 年美国制造业数据拟合。The original Cobb–Douglas paper, fitted to US manufacturing data for 1899–1922.
- Arrow, K. J., Chenery, H. B., Minhas, B. S. & Solow, R. M. (1961). "Capital-Labor Substitution and Economic Efficiency." Review of Economics and Statistics 43(3), 225–250.CES 生产函数的出处(常简称 ACMS)。四位作者里两位后来得了诺奖。The source of the CES production function, usually abbreviated ACMS. Two of the four authors later won the Nobel Prize.
- Shephard, R. W. (1953). Cost and Production Functions.\(\partial C/\partial w=L^c\):与消费者那边的 \(\partial E/\partial p=x^c\) 是同一条引理。\(\partial C/\partial w=L^c\) is the same lemma as \(\partial E/\partial p=x^c\) on the consumer side.
易错点Where it goes wrong
- σ 与 ρ 的关系记成 \(1/(1+\rho)\)。正确是 \(\sigma=\dfrac{1}{1-\rho}\);\(1/(1+\rho)\) 是把 CES 写成 \(\rho\) 在分母那个版本时的形式,两种写法混用必错。Remembering the relation as \(\sigma=1/(1+\rho)\). It is \(\sigma=\dfrac{1}{1-\rho}\); \(1/(1+\rho)\) belongs to the version of CES written with \(\rho\) in the denominator, and mixing the two conventions guarantees an error.
- ρ→0 直接断言等于柯布–道格拉斯。这是 \(1^{\infty}\) 未定式,必须取对数化成 \(0/0\) 再用洛必达。省掉这一步,遇到「证明」类题就丢分。Asserting the ρ→0 limit without proof. It is a \(1^{\infty}\) indeterminate form and needs logs plus LHôpital. Skip the step and any "show that" question is lost.
- 把 σ 理解成「产出对要素的弹性」。σ 是要素比对价格比的弹性,与产出弹性(柯布–道格拉斯里的 \(a,b\))是两回事。Reading σ as an output elasticity. σ is the elasticity of the factor ratio with respect to the price ratio, quite distinct from the output elasticities \(a\) and \(b\) in Cobb–Douglas.
一般均衡与两个福利定理General equilibrium and the two welfare theorems
General equilibrium and the two welfare theorems
把所有市场同时出清写成一个不动点问题——这是「看不见的手」的严格版本。
Writing "all markets clear at once" as a fixed-point problem — the rigorous version of the invisible hand.
核心方程Core equations
展开逐步推导(3 步)Show the 3-step derivation
- 消费者问题The consumer's problem$$x^i\in\arg\max\ u^i(x)\ \text{ s.t. }\ p\cdot x\le p\cdot\omega^i+\textstyle\sum_j\theta^{ij}\pi^j$$注意收入不是外生的:它由禀赋的市场价值加上企业利润的份额内生决定。Income is not exogenous: it is the market value of the endowment plus a share of firm profits, both determined within the model.
- 企业问题The firm's problem$$y^j\in\arg\max\ p\cdot y\ \text{ s.t. }\ y\in Y^j$$\(Y^j\) 是生产可能集。\(Y^j\) is the production possibility set.
- 市场出清Market clearing$$\sum_i x^i=\sum_i\omega^i+\sum_j y^j$$三条缺一条就不是均衡,而不是「写得简略」。这是判断题与论述题的固定扣分点。Drop any one of the three and it is not an equilibrium — not merely a condensed statement of one. Examiners take marks off for this with great regularity.
展开逐步推导(3 步)Show the 3-step derivation
- 瓦尔拉斯定律Walras' law$$z(p)=\textstyle\sum_i x^i(p)-\sum_i\omega^i-\sum_j y^j(p),\qquad p\cdot z(p)=0$$来自每个人的预算约束取等号。后果:一个市场出清 ⇒ 另一个自动出清(两商品情形),所以只能定相对价格,绝对价格水平不定(需要标准化)。It follows from each budget constraint holding with equality. Consequence: if one market clears the other clears automatically (with two goods), so only relative prices are determined and the absolute level needs a normalisation.
- 第一福利定理First welfare theorem$$\text{competitive equilibrium}\ \Longrightarrow\ \text{Pareto optimal}$$只需要偏好局部非饱和。不需要凸性、不需要连续。这是「看不见的手」的严格陈述。All it requires is local non-satiation of preferences — not convexity, not continuity. This is the invisible hand stated rigorously.
- 第二福利定理Second welfare theorem$$\text{Pareto optimal}\ \Longrightarrow\ \text{prices + lump-sum transfers}$$要求凸性(偏好凸、技术凸)。它把「效率」与「公平」分开:先转移禀赋,再让市场干活。This one does require convexity of preferences and technology. It separates efficiency from equity: redistribute endowments first, then let the market work.
交互图Interactive chart
教学算例Worked example
两人两货,\(u^A=u^B=\sqrt{xy}\);A 禀赋 \((10,0)\),B 禀赋 \((0,10)\)。令 \(p_y=1\),求 \(p=p_x\)。
Two agents, two goods, \(u^A=u^B=\sqrt{xy}\); A is endowed with \((10,0)\), B with \((0,10)\). Normalise \(p_y=1\) and solve for \(p=p_x\).
| A 的收入A's income | \(10p\) | 禀赋的市场价值The market value of the endowment |
| A 的需求A's demand | \(x_A=5,\ y_A=5p\) | 份额 ½ ⇒ 各花一半With a share of ½, half of income goes to each good |
| B 的需求B's demand | \(x_B=5/p,\ y_B=5\) | |
| x 市场出清Clearing the x market | \(5+5/p=10\) | 解得 p = 1gives p = 1 |
| y 市场The y market | \(5p+5=10\ \checkmark\) | 自动出清 —— 瓦尔拉斯定律clears automatically — Walras' law |
| 均衡配置Equilibrium allocation | A=(5,5), B=(5,5) | |
| A 的效用增益A's utility gain | 0 → 5 | 交易前 \(\sqrt{10\times0}=0\),交易后 5Before trade \(\sqrt{10\times0}=0\); after trade, 5 |
经典文献Original sources
- Arrow, K. J. & Debreu, G. (1954). "Existence of an Equilibrium for a Competitive Economy." Econometrica 22(3), 265–290.存在性证明,用角谷不动点定理。两位作者分别在 1972、1983 年获诺奖。The existence proof, via Kakutani's fixed-point theorem. The authors received the Nobel Prize in 1972 and 1983 respectively.
- Debreu, G. (1959). Theory of Value. Cowles Foundation Monograph 17.公理化的完整体系,全书不到 100 页,至今是范本。The complete axiomatic treatment, under 100 pages, still a model of exposition.
- Sonnenschein (1972), Mantel (1974), Debreu (1974).SMD 定理:超额需求函数几乎可以是任意形状 ⇒ 一般均衡不保证唯一性与稳定性。这是对该框架最重要的限定。The SMD theorem: excess demand functions can take almost any shape, so general equilibrium guarantees neither uniqueness nor stability. This is the most important qualification the framework carries.
易错点Where it goes wrong
- 均衡只写「供给=需求」。竞争均衡是「一组配置 加 一组价格」,且必须同时满足消费者最优、企业最优、市场出清三条。少写一条就不是均衡的定义。Writing equilibrium as "supply equals demand". A competitive equilibrium is an allocation together with a price vector, satisfying consumer optimisation, firm optimisation and market clearing simultaneously. Omit one and you have not stated the definition.
- 忘了价格只能定到相对水平。瓦尔拉斯定律使方程组少一个独立方程,必须标准化(令某个价格为 1),否则会以为「方程比未知数少一个,无解」。Forgetting that prices are determined only up to scale. Walras' law removes one independent equation, so a normalisation is required (set one price to 1). Without it the system looks underdetermined.
- 把第二福利定理当成「市场自动实现公平」。它要求一次性转移支付先把禀赋挪好,而现实中的转移几乎总是扭曲性的。这是理论与政策之间最大的裂缝。Reading the second welfare theorem as "markets deliver fairness on their own". It requires lump-sum transfers to move endowments first, and real-world transfers are almost always distortionary. This is the widest gap between the theory and any policy built on it.
纳什均衡:古诺与伯特兰Nash equilibrium: Cournot and Bertrand
Nash equilibrium, Cournot and Bertrand competition
每个人对别人的策略做最优反应,且所有人的信念自洽——交点就是均衡。
Everyone best-responds to everyone else and all beliefs are mutually consistent — the intersection is the equilibrium.
核心方程Core equations
展开逐步推导(2 步)Show the 2-step derivation
- 定义最优反应函数Define the best-response function$$BR_i(s_{-i})=\arg\max_{s_i}\pi_i(s_i,s_{-i})$$均衡就是所有最优反应函数的不动点:\(s^*\in BR(s^*)\)。An equilibrium is a fixed point of the best-response correspondence: \(s^*\in BR(s^*)\).
- 存在性Existence$$\text{finite game}\ \Longrightarrow\ \text{mixed-strategy Nash equilibrium exists}$$Nash (1950) 用角谷不动点定理证明——与 Arrow–Debreu 用的是同一件工具。Nash (1950) proved it with Kakutani's fixed-point theorem — the same instrument Arrow and Debreu used.
展开逐步推导(3 步)Show the 3-step derivation
- 写企业 1 的利润Write down firm 1's profit$$\pi_1=\bigl[a-b(q_1+q_2)-c\bigr]q_1=(a-c)q_1-bq_1^2-bq_1q_2$$展开是为了下一步求导时不出错——括号里含 \(q_1\),直接对乘积求导容易漏项。Expanding first avoids slips at the next step: \(q_1\) sits inside the bracket, and differentiating the product directly invites a dropped term.
- 对 q₁ 求导得最优反应Differentiate with respect to q₁ for the best response$$\frac{\partial\pi_1}{\partial q_1}=(a-c)-2bq_1-bq_2=0\ \Longrightarrow\ q_1=\frac{a-c-bq_2}{2b}$$\(-2bq_1\) 里的 2 来自 \(q_1^2\) 求导。这个 2 是古诺与垄断差别的全部来源。The 2 in \(-2bq_1\) comes from differentiating \(q_1^2\). That single 2 is the entire difference between Cournot and monopoly.
- 对称性 ⇒ 令 q₁ = q₂ = qImpose symmetry: q₁ = q₂ = q$$q=\frac{a-c-bq}{2b}\ \Longrightarrow\ 2bq+bq=a-c\ \Longrightarrow\ q^*=\frac{a-c}{3b}$$〔移项:两边同乘 \(2b\),把 \(bq\) 挪到左边合并成 \(3bq\)〕分母上的 3 = n+1,\(n\) 家企业时 \(q^*=\dfrac{a-c}{(n+1)b}\),\(n\to\infty\) 退化到完全竞争。〔Rearranging: multiply through by \(2b\), move \(bq\) to the left and collect into \(3bq\)〕The 3 in the denominator is n+1: with \(n\) firms \(q^*=\dfrac{a-c}{(n+1)b}\), which tends to the competitive outcome as \(n\to\infty\).
展开逐步推导(1 步)Show the 1-step derivation
- 为什么Why$$p_i\gt c\ \Rightarrow\ \text{rival cuts by }\varepsilon\text{ and takes the whole market}$$古诺(选产量)与伯特兰(选价格)给出完全不同的结论——这说明结论对「策略变量是什么」极度敏感,而不是对「有几家企业」敏感。Cournot (quantity setting) and Bertrand (price setting) give completely different answers. The conclusion is acutely sensitive to what the strategic variable is, and hardly at all to how many firms there are.
交互图Interactive chart
教学算例Worked example
线性需求,两家对称,边际成本 10。
Linear demand, two symmetric firms, marginal cost 10.
| 古诺:每家产量Cournot: output per firm | 30 | \((a-c)/3b=90/3\)\((a-c)/3b=90/3\) |
| 古诺:总产量 / 价格Cournot: total output / price | 60 / 40 | |
| 古诺:每家利润Cournot: profit per firm | 900 | \((40-10)\times30\)\((40-10)\times30\) |
| 垄断:产量 / 价格Monopoly: output / price | 45 / 55 | \((a-c)/2b\)\((a-c)/2b\) |
| 垄断:总利润Monopoly: total profit | 2025 | 比两家古诺加总的 1800 更高higher than the 1800 the two Cournot firms make between them |
| 完全竞争:产量 / 价格Perfect competition: output / price | 90 / 10 | \(P=c\)\(P=c\) |
经典文献Original sources
- Cournot, A. A. (1838). Recherches sur les principes mathématiques de la théorie des richesses.比纳什早 112 年就写出了这个均衡——数量竞争的原型。Written 112 years before Nash, and already the prototype of quantity competition.
- Bertrand, J. (1883). "Théorie mathématique de la richesse sociale." Journal des Savants.对古诺的书评,指出改选价格结论就翻转。A review of Cournot pointing out that switching to price competition reverses the conclusion.
- Nash, J. F. (1950). "Equilibrium Points in N-Person Games." PNAS 36(1), 48–49;(1951) Annals of Mathematics 54, 286–295.两页的 PNAS 短文加一篇年鉴长文,奠定了整个非合作博弈论。1994 年诺奖。A two-page note in PNAS and a paper in the Annals founded the whole of non-cooperative game theory. Nobel Prize 1994.
易错点Where it goes wrong
- 把最优反应的一阶条件当成均衡。\(q_1=BR_1(q_2)\) 只是一条曲线;均衡是两条曲线的交点,必须联立。只解一条就报答案是最常见的失分。Mistaking a first-order condition for the equilibrium. \(q_1=BR_1(q_2)\) is only a curve; the equilibrium is where two curves cross, and the pair must be solved jointly. Reporting one curve as the answer is the commonest way marks are lost.
- 混合策略均衡不用「让对手无差异」的条件。求混合策略时,均衡概率由使对手在其纯策略间无差异的条件决定,而不是由自己的收益最大化决定——这个反直觉之处是考试重灾区。Not using the indifference condition for mixed strategies. Equilibrium probabilities are pinned down by making the opponent indifferent across their pure strategies, not by maximising your own payoff. The counter-intuitiveness of this is why it goes wrong so often.
- 以为「企业越多越竞争」是普遍规律。伯特兰里两家就够把价格压到边际成本;古诺里要 \(n\to\infty\)。结论取决于策略变量,不取决于家数。Believing that more firms always means more competition. Two suffice under Bertrand to drive price to marginal cost; under Cournot it takes \(n\to\infty\). The answer depends on the strategic variable, not the headcount.
信息不对称:柠檬市场与信号Asymmetric information: lemons and signalling
Adverse selection, signalling, and screening
当一方比另一方更了解商品质量,市场可能不是定价偏了,而是整个消失。
When one side knows more about quality than the other, the market may not merely misprice — it may disappear.
核心方程Core equations
展开逐步推导(4 步)Show the 4-step derivation
- 给定价格 p,谁愿意卖At a price p, who is willing to sell$$\text{sell}\iff \theta\le p$$保留值低于价格才卖。关键:愿意卖的恰恰是质量差的那批——这就是「逆向」二字的由来。Only those whose reservation value falls below the price. The sellers who come forward are precisely the ones with the worst goods — hence the word "adverse".
- 成交品的平均质量Average quality of what actually trades$$\mathbb{E}[\theta\mid\theta\le p]=\frac{p}{2}$$均匀分布截断后的均值是区间中点。Truncating a uniform distribution leaves the midpoint of the remaining interval.
- 买方的支付意愿The buyer's willingness to pay$$\text{WTP}=k\cdot\frac{p}{2}$$买方理性预期到自己买到的是次品,据此出价。The buyer rationally anticipates receiving a poor unit and bids accordingly.
- 市场存活的条件The condition for the market to survive$$k\cdot\frac{p}{2}\ge p\iff k\ge 2$$即使 \(k\gt1\)(每一笔交易都有得益),只要 \(k\lt2\) 市场就全面崩溃。不是萎缩,是唯一均衡为零交易。Even with \(k\gt1\), so that every single trade creates surplus, any \(k\lt2\) collapses the market entirely. Not shrinks it — the unique equilibrium is zero trade.
展开逐步推导(3 步)Show the 3-step derivation
- 单交叉条件是信号有效的前提Single crossing is what makes a signal work$$\frac{\partial}{\partial\theta}\Bigl(\frac{\partial c(e,\theta)/\partial e}{\partial u/\partial w}\Bigr)\lt0$$高能力者获取信号的边际成本更低。没有这一条,任何信号都无法分离类型——教育之所以能当信号,靠的是「学得动」而不是「学到了什么」。The high type must find the signal cheaper at the margin. Without that condition no signal can separate types — education works as a signal because the able find it easier, not because of anything they learn.
- 分离均衡Separating equilibrium$$e^*:\ w_H-c(e^*,\theta_L)\le w_L$$分离均衡通常有连续多个(不唯一),且社会性浪费:教育本身不提高生产率,纯粹烧钱证明自己。Separating equilibria are typically a continuum rather than unique, and they are socially wasteful: education raises no one\'s productivity here and is pure expenditure on proving a point.
- Rothschild–Stiglitz 的负面结论The negative result of Rothschild–Stiglitz$$\text{no pooling equilibrium; separating may fail to exist}$$竞争性保险市场可能根本没有均衡——这是对「市场总能出清」最锋利的反例。A competitive insurance market may have no equilibrium at all — the sharpest available counterexample to the presumption that markets clear.
交互图Interactive chart
教学算例Worked example
质量 \(\theta\sim U[0,1]\),卖方保留值 \(\theta\),买方估值 \(k\theta\)。
Quality \(\theta\sim U[0,1]\), seller reservation value \(\theta\), buyer valuation \(k\theta\).
| k = 1.5k = 1.5 | WTP = 0.75p < p | 市场崩溃——尽管每笔交易都创造 50% 的剩余Market collapses — even though every trade creates 50% surplus |
| k = 2.0k = 2.0 | WTP = 1.00p = p | 临界:任意 p 都是均衡(退化)Knife-edge: any p is an equilibrium (degenerate) |
| k = 2.5k = 2.5 | WTP = 1.25p > p | 市场存活,全部质量都能成交Market survives; every quality level trades |
经典文献Original sources
- Akerlof, G. A. (1970). "The Market for Lemons: Quality Uncertainty and the Market Mechanism." Quarterly Journal of Economics 84(3), 488–500.被数家顶刊拒稿后才发表,理由包括「太琐碎」。2001 年诺奖。Rejected by several leading journals before publication, on grounds that included triviality. Nobel Prize 2001.
- Spence, M. (1973). "Job Market Signaling." Quarterly Journal of Economics 87(3), 355–374.教育作为信号:即使教育完全不提高生产率,高能力者仍愿意购买它。Education as a signal: the able buy it even when it raises productivity not at all.
- Rothschild, M. & Stiglitz, J. (1976). "Equilibrium in Competitive Insurance Markets." QJE 90(4), 629–649.筛选(screening):不知情方设计合同菜单让对方自选。给出了均衡可能不存在的著名结果。Screening: the uninformed side designs a menu of contracts and lets the informed side sort itself. Contains the celebrated non-existence result.
- Holmström, B. (1979). "Moral Hazard and Observability." Bell Journal of Economics 10(1), 74–91.道德风险与充分统计量原理——逆向选择之外的另一半信息经济学。2016 年诺奖。Moral hazard and the sufficient statistic result — the other half of information economics. Nobel Prize 2016.
易错点Where it goes wrong
- 把逆向选择与道德风险混为一谈。逆向选择是签约前的隐藏信息(不知道对方是什么类型);道德风险是签约后的隐藏行动(看不见对方做了什么)。对策完全不同:前者靠信号/筛选,后者靠激励合同。Conflating adverse selection with moral hazard. Adverse selection is hidden information before contracting (you do not know the type); moral hazard is hidden action after contracting (you cannot see what they do). The remedies differ entirely: signalling and screening for the first, incentive contracts for the second.
- 以为「有交易得益市场就会存在」。柠檬模型的核心正是反例:\(k=1.5\) 时每笔交易都有 50% 的剩余,市场却完全消失。存在得益 ≠ 存在均衡。Assuming that gains from trade imply a market. The lemons model is the standing counterexample: at \(k=1.5\) every trade would create 50% surplus and the market vanishes anyway. Gains from trade ≠ existence of equilibrium.
- 忘记买方是理性预期的。常见错误是让买方按平均质量 \(\mathbb{E}[\theta]=0.5\) 出价。不对——买方知道只有 \(\theta\le p\) 的会卖,所以按截断后的条件均值 \(p/2\) 出价。这一步是整个模型的引擎。Forgetting that the buyer has rational expectations. The usual slip is to let the buyer bid on unconditional average quality \(\mathbb{E}[\theta]=0.5\). No — the buyer knows only \(\theta\le p\) is offered and bids on the truncated conditional mean \(p/2\). That step is the engine of the whole model.
Solow 增长模型The Solow growth model
The Solow–Swan growth model
资本积累有收益递减,所以长期增长必须来自技术——这是增长论的出发点与自我否定。
Capital accumulation runs into diminishing returns, so long-run growth must come from technology — the starting point of growth theory and its own refutation.
核心方程Core equations
展开逐步推导(2 步)Show the 2-step derivation
- 从总量到有效人均From aggregates to effective labour units$$k\equiv\frac{K}{AL},\qquad \frac{\dot k}{k}=\frac{\dot K}{K}-n-g$$口径必须钉死:\(k\) 是「有效劳动人均资本」,不是人均资本,也不是总资本。差一个口径,后面的 \(\delta+n+g\) 就写不对。Fix the units before anything else: \(k\) is capital per unit of effective labour, not per worker and not the aggregate. Get the units wrong and \(\delta+n+g\) comes out wrong.
- 写出净积累Write net accumulation$$\dot k=\underbrace{sf(k)}_{\text{investment}}-\underbrace{(\delta+n+g)k}_{\text{break-even}}$$\(\delta\) 折旧、\(n\) 人口稀释、\(g\) 技术稀释——三者都是「要维持 k 不变必须补上的量」,所以并列相加。Depreciation \(\delta\), population dilution \(n\) and technological dilution \(g\) are all amounts that must be replaced merely to hold k constant, which is why they simply add.
展开逐步推导(3 步)Show the 3-step derivation
- 令 k̇ = 0 解稳态Set k̇ = 0 and solve$$sk^{\alpha}=(\delta+n+g)k\ \Longrightarrow\ k^{1-\alpha}=\frac{s}{\delta+n+g}$$〔两边同除 \(k^{\alpha}\),指数相减:\(k^{1}/k^{\alpha}=k^{1-\alpha}\)〕再取 \(\frac{1}{1-\alpha}\) 次幂。〔Divide both sides by \(k^{\alpha}\) and subtract exponents: \(k^{1}/k^{\alpha}=k^{1-\alpha}\)〕then raise to the power \(\frac{1}{1-\alpha}\).
- 黄金律:最大化稳态消费Golden rule: maximise steady-state consumption$$c^*=f(k^*)-(\delta+n+g)k^*,\qquad \frac{dc^*}{dk^*}=f'(k^*)-(\delta+n+g)=0$$对柯布–道格拉斯,\(f'(k)=\alpha k^{\alpha-1}\),代入得 黄金律储蓄率就等于资本份额 α。这是一个漂亮到不像话的结果。For Cobb–Douglas, \(f'(k)=\alpha k^{\alpha-1}\), which gives a golden-rule saving rate equal to the capital share α — a result almost too neat to be true.
- 收敛速度Speed of convergence$$\dot k\approx-(1-\alpha)(\delta+n+g)(k-k^*)$$缺口每年缩小 \((1-\alpha)(\delta+n+g)\)。这正是实证「条件收敛速度约 2%/年」的理论对应物。The gap closes at \((1-\alpha)(\delta+n+g)\) a year. This is the theoretical counterpart of the empirical finding that conditional convergence runs at roughly 2% a year.
交互图Interactive chart
教学算例Worked example
持平投资率 \(\delta+n+g=0.08\)。
Break-even investment rate \(\delta+n+g=0.08\).
| 稳态资本 \(k^*\)Steady-state capital \(k^*\) | 3.9528 | \((0.2/0.08)^{1.5}\)\((0.2/0.08)^{1.5}\) |
| 稳态产出 \(y^*\)Steady-state output \(y^*\) | 1.5811 | \(=k^{*1/3}\)\(=k^{*1/3}\) |
| 稳态消费 \(c^*\)Steady-state consumption \(c^*\) | 1.2649 | \(=(1-s)y^*\)\(=(1-s)y^*\) |
| 黄金律资本Golden-rule capital | 8.5052 | 口算 8.5046 是错的——脚本算出 8.50517The mental arithmetic gave 8.5046 and was wrong; the script returns 8.50517 |
| 黄金律储蓄率Golden-rule saving rate | 1/3 = 0.3333 | 等于资本份额 αEqual to the capital share α |
| 黄金律消费Golden-rule consumption | 1.3608 | 比当前 c* 高 7.6%7.6% above the current c* |
| 收敛速度Speed of convergence | 5.33%/年 | \((1-\alpha)(\delta+n+g)\)\((1-\alpha)(\delta+n+g)\) |
| 缺口半衰期Half-life of the gap | 13.0 年 | \(\ln2/0.0533\)\(\ln2/0.0533\) |
经典文献Original sources
- Solow, R. M. (1956). "A Contribution to the Theory of Economic Growth." Quarterly Journal of Economics 70(1), 65–94.1987 年诺奖。Nobel Prize 1987.
- Swan, T. W. (1956). "Economic Growth and Capital Accumulation." Economic Record 32(2), 334–361.同年独立提出,故正式称 Solow–Swan 模型。Published independently the same year, which is why the model is properly Solow–Swan.
- Phelps, E. S. (1961). "The Golden Rule of Accumulation: A Fable for Growthmen." American Economic Review 51(4), 638–643.黄金律。2006 年诺奖。The golden rule. Nobel Prize 2006.
- Mankiw, N. G., Romer, D. & Weil, D. N. (1992). "A Contribution to the Empirics of Economic Growth." QJE 107(2), 407–437.加入人力资本的扩展 Solow,实证上把跨国收入差异解释力提到约 80%。Solow augmented with human capital; empirically it raises the share of cross-country income differences explained to around 80%.
易错点Where it goes wrong
- 把 k 的口径搞混。有效劳动人均 \(K/(AL)\)、人均 \(K/L\)、总量 \(K\) 三者的积累方程分母不同:依次是 \(\delta+n+g\)、\(\delta+n\)、\(\delta\)。写方程前先声明口径,这是本模型最高频的错。Losing track of the units of k. Per effective worker \(K/(AL)\), per worker \(K/L\) and aggregate \(K\) carry different accumulation terms: \(\delta+n+g\), \(\delta+n\) and \(\delta\) respectively. State the units before writing the equation — this is the single most frequent error in the model.
- 以为提高储蓄率能提高长期增长率。提高 \(s\) 只提高水平 \(k^*,y^*\),长期增长率恒等于外生的 \(g\)。这是 Solow 模型最反直觉、也最重要的结论,模型 8 正是为解决这一点而生。Believing a higher saving rate raises the long-run growth rate. Raising \(s\) raises the levels \(k^*\) and \(y^*\); the long-run growth rate is exogenous \(g\) and nothing else. This is the most counter-intuitive and most important result in the model, and Model 8 exists to answer it.
- 认为储蓄越多越好。\(s\gt\alpha\) 时经济动态无效:资本过多,减少储蓄可以让每一代人的消费都上升。Assuming more saving is always better. When \(s\gt\alpha\) the economy is dynamically inefficient: there is too much capital, and cutting saving would raise consumption for every generation.
Ramsey–Cass–Koopmans 最优增长Ramsey–Cass–Koopmans optimal growth
Optimal growth: the Ramsey–Cass–Koopmans model
储蓄率不再是天上掉下来的参数,而是家庭在无限期上权衡出来的。
The saving rate stops being a parameter handed down from above and becomes something households trade off over an infinite horizon.
核心方程Core equations
展开逐步推导(4 步)Show the 4-step derivation
- 构造现值汉密尔顿函数Form the present-value Hamiltonian$$H=\frac{c^{1-\theta}-1}{1-\theta}+\mu\bigl[f(k)-c-(\delta+n)k\bigr]$$\(\mu\) 是资本的影子价格。控制变量是 \(c\),状态变量是 \(k\)。\(\mu\) is the shadow price of capital. The control is \(c\), the state is \(k\).
- 一阶条件与协态方程First-order and costate conditions$$\frac{\partial H}{\partial c}=c^{-\theta}-\mu=0;\qquad \dot\mu=\rho\mu-\frac{\partial H}{\partial k}$$第一条把 \(\mu\) 与消费的边际效用绑定;第二条是资产定价方程。The first ties \(\mu\) to the marginal utility of consumption; the second is an asset pricing equation.
- 消掉 μ 得欧拉方程Eliminate μ to get the Euler equation$$\frac{\dot c}{c}=\frac{1}{\theta}\bigl[f'(k)-\delta-\rho\bigr]$$对 \(c^{-\theta}=\mu\) 取对数再求导:\(-\theta\dfrac{\dot c}{c}=\dfrac{\dot\mu}{\mu}\)。1/θ 是跨期替代弹性:资本回报超过贴现率多少,消费就以多快的速度增长。Take logs of \(c^{-\theta}=\mu\) and differentiate: \(-\theta\dfrac{\dot c}{c}=\dfrac{\dot\mu}{\mu}\). 1/θ is the elasticity of intertemporal substitution: it converts the excess of the return on capital over the discount rate into a rate of consumption growth.
- 横截性条件Transversality condition$$\lim_{t\to\infty}e^{-\rho t}\mu(t)k(t)=0$$这一条最常被遗忘。没有它,欧拉方程有无穷多条解路径(可以永远借下去),鞍点路径的唯一性正是靠它钉住的。The most frequently forgotten condition in the model. Without it the Euler equation admits infinitely many paths (one may borrow for ever); it is what pins the solution down to the unique saddle path.
展开逐步推导(2 步)Show the 2-step derivation
- 令 ċ = 0Set ċ = 0$$f'(k^*)=\rho+\delta\ \Longrightarrow\ k^*=\Bigl(\frac{\alpha}{\rho+\delta}\Bigr)^{1/(1-\alpha)}$$贴现率 \(\rho\) 进了分母,ρ 越大稳态资本越少——不耐心的社会积累得少。The discount rate \(\rho\) enters the denominator, so a larger ρ means a smaller steady-state capital stock: an impatient society accumulates less.
- 与黄金律比较Compare with the golden rule$$\rho\gt n+g\ \Longrightarrow\ k^*\lt k_{\text{gold}}$$Ramsey 经济永远不会动态无效。这是与 Solow 最重要的差别:Solow 的 \(s\) 可以随便设得过高,Ramsey 里理性家庭绝不会那么做。A Ramsey economy is never dynamically inefficient. That is the important contrast with Solow: an exogenous \(s\) can be set too high, whereas optimising households never choose to.
交互图Interactive chart
教学算例Worked example
与模型 6 用同一个生产函数,便于直接对照。
Same production function as Model 6, so the two are directly comparable.
| 稳态 \(k^*\)Steady state \(k^*\) | 7.1278 | \((\alpha/(\rho+\delta))^{1.5}\)\((\alpha/(\rho+\delta))^{1.5}\) |
| \(f'(k^*)\)\(f'(k^*)\) | 0.09 | \(=\rho+\delta\) ✓\(=\rho+\delta\) ✓ |
| 黄金律 \(k_{gold}\)(模型 6)Golden-rule \(k_{gold}\) (Model 6) | 8.5052 | |
| 贴现造成的缺口Gap created by discounting | 1.3774 | 理性家庭主动选择低于黄金律的资本Optimising households deliberately choose less capital than the golden rule |
| k=5 处的消费增长率Consumption growth at k=5 | 1.20%/年 | \((f'(5)-\delta-\rho)/\theta\),为正 ⇒ 仍在积累\((f'(5)-\delta-\rho)/\theta\), positive, so accumulation continues |
经典文献Original sources
- Ramsey, F. P. (1928). "A Mathematical Theory of Saving." Economic Journal 38(152), 543–559.26 岁写就,两年后去世。Keynes 称之为「数学经济学有史以来最杰出的贡献之一」。Written at 26; the author died two years later. Keynes called it one of the most remarkable contributions ever made to mathematical economics.
- Cass, D. (1965). "Optimum Growth in an Aggregative Model of Capital Accumulation." Review of Economic Studies 32(3), 233–240.把 Ramsey 问题嵌入新古典增长框架。Embedded Ramsey's problem in the neoclassical growth framework.
- Koopmans, T. C. (1965). "On the Concept of Optimal Economic Growth."与 Cass 同年独立完成,故三人并称。1975 年诺奖。Completed independently in the same year, hence the three names. Nobel Prize 1975.
易错点Where it goes wrong
- 忘掉横截性条件。只有欧拉方程时,任何一条初始消费都能生成一条满足欧拉方程的路径;横截性条件把它们筛到唯一一条鞍点路径上。论述题里漏掉这一条几乎必然扣分。Omitting the transversality condition. With only the Euler equation, any initial consumption generates a path satisfying it; the transversality condition is what selects the unique saddle path. Leaving it out of a written answer costs marks almost every time.
- 把 1/θ 说成风险规避系数。\(\theta\) 是相对风险规避系数,\(1/\theta\) 是跨期替代弹性。在 CRRA 效用下二者互为倒数,这是一个受批评的巧合(Epstein–Zin 偏好就是为拆开它们而生)。Calling 1/θ the coefficient of risk aversion. \(\theta\) is relative risk aversion; \(1/\theta\) is the elasticity of intertemporal substitution. Under CRRA they are reciprocals, a coincidence that has drawn steady criticism — Epstein–Zin preferences exist precisely to prise them apart.
- 以为稳态一定优于黄金律。\(k^*\lt k_{\text{gold}}\) 意味着稳态消费低于黄金律消费——但这是最优的,因为达到黄金律要求当下牺牲太多。「消费更高」不等于「更好」。Assuming the steady state must beat the golden rule. \(k^*\lt k_{\text{gold}}\) means steady-state consumption is lower than at the golden rule — and that is optimal, because reaching the golden rule would cost too much today. Higher consumption is not the same thing as better.
内生增长与创造性破坏Endogenous growth and creative destruction
Endogenous growth: AK, Romer, and Schumpeterian models
增长率不再外生:它由研发的私人回报决定,而回报又被下一次创新所摧毁。
The growth rate is no longer exogenous: it is set by the private return to research, and that return is destroyed by the next innovation.
核心方程Core equations
展开逐步推导(1 步)Show the 1-step derivation
- 为什么没有收敛Why there is no convergence$$\frac{\dot K}{K}=s A-\delta\quad(\text{independent of }K)$$\(f'(k)=A\) 是常数,收益不再递减,所以储蓄率永久影响增长率。代价是这个结论对「α 恰好等于 1」极度敏感——稍微小一点就退回 Solow。\(f'(k)=A\) is constant, so returns no longer diminish and the saving rate permanently affects the growth rate. The cost is that the result is acutely sensitive to α being exactly 1; anything less and the model reverts to Solow.
展开逐步推导(4 步)Show the 4-step derivation
- 创新的到达与幅度Arrival rate and size of innovations$$\text{Poisson rate}=\lambda n,\qquad \text{productivity}\times\gamma\ (\gamma\gt1)$$\(n\) 是投入研发的人数。单位时间的对数产出增长 = 到达率 × 每次的对数跳幅。\(n\) is the number of researchers. Log output growth per unit time is the arrival rate times the log size of each jump.
- 增长率The growth rate$$g=\lambda n\cdot\ln\gamma$$取对数是因为增长是乘法的:\(\ln(\gamma^m)=m\ln\gamma\)。Logs appear because growth is multiplicative: \(\ln(\gamma^m)=m\ln\gamma\).
- 创新的价值——关键一步The value of an innovation — the crucial step$$V=\frac{\pi}{r+\lambda n}$$分母里的 λn 就是创造性破坏:现任垄断者以 \(\lambda n\) 的速率被下一个创新者取代,所以它的租金必须按 \(r+\lambda n\) 而不是 \(r\) 来贴现。研发越热,每项创新越不值钱。The λn in the denominator is creative destruction. The incumbent monopolist is displaced at rate \(\lambda n\), so its rents must be discounted at \(r+\lambda n\) rather than \(r\). The hotter the research race, the less any single innovation is worth.
- 研究套利条件Research arbitrage$$w=\lambda V\ \Longrightarrow\ n^*$$研发者的工资等于其边际产出(创新到达率 × 创新价值)。这条方程与劳动市场出清一起,把 \(n^*\) 定下来,从而定住增长率。A researcher's wage equals their marginal product (arrival rate times the value of an innovation). Together with labour market clearing this pins down \(n^*\), and with it the growth rate.
交互图Interactive chart
教学算例Worked example
到达率 \(\lambda n=0.2\)/年。
Arrival rate \(\lambda n=0.2\) per year.
| 增长率 gGrowth rate g | 8.11% | \(0.1\times2\times\ln1.5\)\(0.1\times2\times\ln1.5\) |
| 创新价值 VValue of an innovation V | 4.00 | \(\pi/(r+\lambda n)=1/0.25\)\(\pi/(r+\lambda n)=1/0.25\) |
| 若无创造性破坏Value absent creative destruction | 20.00 | \(\pi/r=1/0.05\)\(\pi/r=1/0.05\) |
| 破坏造成的价值折损Value lost to destruction | 80% | 4 / 204 / 20 |
| AK 对照:A=0.5, s=0.2, δ=0.05AK for comparison: A=0.5, s=0.2, δ=0.05 | g = 5% | \(sA-\delta\)\(sA-\delta\) |
经典文献Original sources
- Romer, P. M. (1990). "Endogenous Technological Change." Journal of Political Economy 98(5), S71–S102.知识的非竞争性 + 垄断竞争,水平创新(品种增加)。2018 年诺奖。Non-rivalry of knowledge plus monopolistic competition; horizontal innovation (more varieties). Nobel Prize 2018.
- Aghion, P. & Howitt, P. (1992). "A Model of Growth Through Creative Destruction." Econometrica 60(2), 323–351.垂直创新(质量阶梯)。1987 年投稿,历时五年才发表。2025 年诺贝尔经济学奖(与 Joel Mokyr 共享,表彰「解释创新驱动的经济增长」)。Vertical innovation (quality ladders). Submitted in 1987 and five years in press. Nobel Prize in Economic Sciences 2025, shared with Joel Mokyr, for explaining innovation-driven economic growth.
- Grossman, G. M. & Helpman, E. (1991). Innovation and Growth in the Global Economy.把内生增长嵌入开放经济与贸易。Endogenous growth embedded in open economies and trade.
- Jones, C. I. (1995). "R&D-Based Models of Economic Growth." JPE 103(4), 759–784.尖锐的实证批评:研发人员数十年间增长数十倍,增长率却没变 ⇒ 一代内生增长模型的「规模效应」被数据否定。A sharp empirical objection: research employment rose many times over across decades while growth rates did not, which refutes the scale effect built into the first generation of models.
易错点Where it goes wrong
- 把创新价值写成 π/r。分母必须是 \(r+\lambda n\)。漏掉 \(\lambda n\) 等于假设垄断租金永远持续,那样就没有「创造性破坏」了——这是整个模型的名字所在。Writing the value of an innovation as π/r. The denominator must be \(r+\lambda n\). Dropping \(\lambda n\) assumes monopoly rents last for ever, which removes the creative destruction the model is named after.
- 认为内生增长理论必然支持补贴研发。模型里有两个方向相反的外部性:知识溢出(研发不足)与商业窃取(研发过度)。净效应取决于参数,不是定论。Assuming the theory implies research subsidies. Two externalities point in opposite directions: knowledge spillovers (too little research) and business stealing (too much). The net effect depends on parameters and is not settled.
- 忽视 Jones (1995) 的批评。一代内生增长模型隐含「研发人数翻倍则增长率翻倍」的规模效应,与二战后数据严重冲突。半内生增长模型正是为修补这一点而生。Ignoring Jones (1995). The first generation of models implies that doubling research employment doubles the growth rate, which post-war data contradict outright. Semi-endogenous growth models exist to repair exactly this.
实际经济周期(RBC)Real business cycles
Real business cycles
把 Ramsey 模型加上随机技术冲击,波动就成了最优反应而非市场失灵。
Add stochastic technology shocks to the Ramsey model and fluctuations become an optimal response rather than a market failure.
核心方程Core equations
展开逐步推导(3 步)Show the 3-step derivation
- 写欧拉方程(离散时间)Write the Euler equation in discrete time$$\frac{1}{c_t}=\beta\,\mathbb{E}_t\Bigl[\frac{1}{c_{t+1}}\cdot\alpha z_{t+1}k_{t+1}^{\alpha-1}\Bigr]$$与模型 7 的连续时间欧拉方程是同一条式子的离散版:今天少吃一口的边际损失 = 明天多吃的贴现边际收益。The same equation as the continuous-time Euler equation of Model 7: the marginal loss from eating one less unit today equals the discounted marginal gain tomorrow.
- 猜解:储蓄率为常数Guess a constant saving rate$$k_{t+1}=\sigma y_t\ \Longrightarrow\ c_t=(1-\sigma)y_t$$对数效用 + 全折旧 + 柯布–道格拉斯这三条同时成立时,收入效应与替代效应正好抵消,储蓄率为常数。With log utility, full depreciation and Cobb–Douglas holding together, income and substitution effects cancel exactly and the saving rate is constant.
- 代回欧拉方程定 σSubstitute back to pin down σ$$\frac{1}{(1-\sigma)y_t}=\frac{\alpha\beta}{(1-\sigma)\sigma y_t} \ \Longrightarrow\ \sigma=\alpha\beta$$〔约分:\(y_{t+1}\) 上下相消,\(k_{t+1}=\sigma y_t\) 代入分母〕储蓄率 αβ 与冲击无关,所以才有闭式解。〔Cancelling: \(y_{t+1}\) divides out and \(k_{t+1}=\sigma y_t\) goes into the denominator〕The saving rate αβ is independent of the shock, which is what makes the closed form possible.
展开逐步推导(2 步)Show the 2-step derivation
- 当期产出弹性Impact elasticity of output$$\frac{\partial\ln y_t}{\partial\ln z_t}=1$$\(k_t\) 是前定变量(上一期决定的),所以 TFP 冲击当期一比一传到产出。\(k_t\) is predetermined (it was set last period), so a TFP shock passes one-for-one into output on impact.
- 内在传导:资本积累Internal propagation through capital$$\hat k_{t+1}=\hat y_t=\hat z_t+\alpha\hat k_t$$这就是 RBC 的传导机制——冲击本身是 AR(1),但资本积累把它拉长成更持久的产出波动。This is the RBC propagation mechanism: the shock itself is an AR(1), and capital accumulation stretches it into a more persistent path for output.
交互图Interactive chart
教学算例Worked example
季度校准,标准 RBC 参数。
Quarterly calibration, standard RBC parameters.
| 储蓄率 αβSaving rate αβ | 0.3564 | 常数,不随冲击变化Constant, invariant to the shock |
| 消费份额Consumption share | 0.6436 | |
| TFP 路径 t=0…5TFP path, t=0…5 | 1.000, 0.950, 0.903, 0.857, 0.815, 0.774 | ρ_z 的幂Powers of ρ_z |
| 产出当期弹性Impact elasticity of output | 1.00 | k 前定 ⇒ 一比一k predetermined, hence one-for-one |
| TFP 半衰期TFP half-life | 13.5 季度 | \(\ln0.5/\ln0.95\)\(\ln0.5/\ln0.95\) |
经典文献Original sources
- Kydland, F. E. & Prescott, E. C. (1982). "Time to Build and Aggregate Fluctuations." Econometrica 50(6), 1345–1370.RBC 的开山之作,「校准」方法论也由此确立。2004 年诺奖。The founding paper of RBC, and the origin of the calibration methodology. Nobel Prize 2004.
- Long, J. B. & Plosser, C. I. (1983). "Real Business Cycles." Journal of Political Economy 91(1), 39–69.给出了上面那个闭式可解的特例,是理解 RBC 机制最好的入口。Contains the closed-form special case above, which remains the best way into the mechanism.
- Summers, L. H. (1986). "Some Skeptical Observations on Real Business Cycle Theory." Minneapolis Fed Quarterly Review.最著名的批评:Solow 余值究竟是技术冲击,还是要素利用率的测量误差?The best-known criticism: is the Solow residual a technology shock, or mismeasured factor utilisation?
- Smets, F. & Wouters, R. (2007). "Shocks and Frictions in US Business Cycles." AER 97(3), 586–606.RBC 的现代后裔:加上价格粘性、工资粘性、习惯形成后贝叶斯估计,成为各国央行的主力模型。The modern descendant: add sticky prices, sticky wages and habit formation, estimate by Bayesian methods, and the result is the workhorse model of central banks.
易错点Where it goes wrong
- 把「校准」当成「估计」。校准是从微观证据或长期均值挑参数值,不做统计推断;它的辩护理由是「模型必错,估计出的标准误没有意义」。这是方法论立场,不是技术细节。Treating calibration as estimation. Calibration takes parameter values from microeconomic evidence or long-run averages and performs no inference; its defence is that since the model is certainly false, standard errors around its estimates mean little. This is a methodological position, not a technical detail.
- 以为 RBC 说「衰退是好事」。模型说的是在给定冲击下,观察到的波动是最优反应,因此稳定政策无益。它不说冲击本身是好事。这个区分在论述题里价值很高。Reading RBC as the claim that recessions are good. The claim is that given the shocks, observed fluctuations are an optimal response, and therefore stabilisation policy does not help. It says nothing about the shocks themselves being desirable. The distinction is worth marks in any written answer.
- 把 Solow 余值直接当技术冲击。余值里混着要素利用率、加成率、测量误差。Summers (1986) 与 Basu–Fernald–Kimball (2006) 的修正表明,纠正后的技术冲击效应与原始 RBC 结论方向相反。Taking the Solow residual as a technology shock. The residual mixes in factor utilisation, markups and measurement error. The corrections in Summers (1986) and Basu–Fernald–Kimball (2006) reverse the sign of the estimated response.
新凯恩斯三方程The three-equation New Keynesian model
The three-equation New Keynesian model
一条动态 IS、一条菲利普斯曲线、一条利率规则——现代央行的思维骨架。
A dynamic IS curve, a Phillips curve and an interest rate rule — the skeleton of how central banks now think.
核心方程Core equations
展开逐步推导(3 步)Show the 3-step derivation
- IS 曲线就是欧拉方程The IS curve is the Euler equation$$\frac{1}{c_t}=\beta\mathbb{E}_t\Bigl[\frac{1+i_t}{1+\pi_{t+1}}\frac{1}{c_{t+1}}\Bigr]$$与模型 7、9 是同一条欧拉方程,只是写成了缺口形式。\(1/\sigma\) 又是跨期替代弹性——实际利率越高,越愿意把消费推后。The same Euler equation as Models 7 and 9, written in gap form. Once again \(1/\sigma\) is the elasticity of intertemporal substitution: the higher the real rate, the more willing the household is to postpone consumption.
- 菲利普斯曲线来自 Calvo 定价The Phillips curve comes from Calvo pricing$$\kappa=\frac{(1-\theta)(1-\beta\theta)}{\theta}(\sigma+\varphi)$$\(\theta\) 是每期不能调价的概率。\(\theta\to0\)(价格完全灵活)时 \(\kappa\to\infty\),曲线变垂直,货币重归中性——新凯恩斯模型退化成 RBC。\(\theta\) is the probability of not being able to reset the price. As \(\theta\to0\) (fully flexible prices) \(\kappa\to\infty\), the curve becomes vertical and money is neutral again — the New Keynesian model collapses back into RBC.
- 泰勒原则The Taylor principle$$\phi_{\pi}\gt1$$名义利率对通胀的反应必须大于一比一,实际利率才会上升、才能压住通胀。\(\phi_{\pi}\lt1\) 时均衡不确定,会出现自我实现的通胀波动——这被广泛用来解释 1970 年代的大通胀。The nominal rate must respond more than one-for-one to inflation, otherwise the real rate does not rise and inflation is not restrained. With \(\phi_{\pi}\lt1\) the equilibrium is indeterminate and self-fulfilling inflation becomes possible — the standard account of the Great Inflation of the 1970s.
交互图Interactive chart
教学算例Worked example
\(i=2+\pi+0.5(\pi-2)+0.5\cdot\text{gap}\),即 \(r^*=2\%\)、\(\pi^*=2\%\)、\(\phi_{\pi}=1.5\)、\(\phi_x=0.5\)。
\(i=2+\pi+0.5(\pi-2)+0.5\cdot\text{gap}\), that is \(r^*=2\%\), \(\pi^*=2\%\), \(\phi_{\pi}=1.5\), \(\phi_x=0.5\).
| π=2%, gap=0π=2%, gap=0 | i = 4.0% | 通胀达标、产出达标 ⇒ 中性名义利率 = r*+π*Inflation and output both on target, so the neutral nominal rate is r*+π* |
| π=4%, gap=+1%π=4%, gap=+1% | i = 7.5% | 通胀超标 2pp ⇒ 名义利率多加 3ppInflation 2pp above target draws a 3pp rise in the nominal rate |
| π=1%, gap=−2%π=1%, gap=−2% | i = 1.5% | 通缩 + 衰退 ⇒ 大幅宽松Disinflation plus recession calls for substantial easing |
| 实际利率对通胀的斜率Slope of the real rate in inflation | +0.5 | \(\phi_{\pi}-1=0.5\gt0\) ⇒ 满足泰勒原则\(\phi_{\pi}-1=0.5\gt0\), so the Taylor principle holds |
经典文献Original sources
- Calvo, G. A. (1983). "Staggered Prices in a Utility-Maximizing Framework." Journal of Monetary Economics 12(3), 383–398.「每期以固定概率能调价」的定价假设,因为可解性极好而成为业界标准。The assumption that a firm may reset its price with fixed probability each period. Its tractability made it the industry standard.
- Taylor, J. B. (1993). "Discretion versus Policy Rules in Practice." Carnegie-Rochester Conference Series on Public Policy 39, 195–214.泰勒规则。原意是描述 1987–92 年美联储的实际行为,后来变成规范性基准。The Taylor rule. Intended as a description of what the Federal Reserve actually did between 1987 and 1992, later adopted as a normative benchmark.
- Clarida, R., Galí, J. & Gertler, M. (1999). "The Science of Monetary Policy: A New Keynesian Perspective." Journal of Economic Literature 37(4), 1661–1707.把三方程整理成今天教科书的形式,并给出最优政策分析。Assembled the three equations into their textbook form and worked out optimal policy.
- Galí, J. (2015). Monetary Policy, Inflation, and the Business Cycle (2nd ed.). Princeton University Press.标准研究生教材,三方程的完整微观基础推导都在这里。The standard graduate text; the full microfoundations of the three equations are here.
易错点Where it goes wrong
- 把 NKPC 当成传统菲利普斯曲线。传统版本是「通胀 vs 失业」的可利用的权衡;新凯恩斯版本是前瞻的:\(\pi_t\) 取决于预期未来通胀,不存在长期权衡,而且理论上「反通胀可以无成本」(只要央行可信)——这与经验事实的紧张是该模型的老问题。Reading the NKPC as the traditional Phillips curve. The traditional version offers an exploitable trade-off between inflation and unemployment; the New Keynesian version is forward-looking: \(\pi_t\) depends on expected future inflation, there is no long-run trade-off, and in principle disinflation is costless given a credible central bank — a prediction whose tension with the evidence is a long-standing problem for the model.
- 忘记泰勒规则里的利率是名义利率。判断政策松紧要看实际利率 \(i-\mathbb{E}\pi\)。名义利率上升但通胀上升更多,实际是在放松。这是新闻评论里最常见的错误。Forgetting that the rate in the Taylor rule is nominal. Whether policy is tight or loose depends on the real rate \(i-\mathbb{E}\pi\). A nominal rate that rises by less than inflation is an easing. This is the commonest error in press commentary.
- 以为 φ_π 大于 1 就万事大吉。在零利率下限(ZLB)处泰勒规则无法执行,不确定性重新出现——这正是 2009 年后前瞻指引与量化宽松的理论动机。Assuming φ_π > 1 settles the matter. At the zero lower bound the rule cannot be implemented and indeterminacy returns — which is precisely the motivation for forward guidance and quantitative easing after 2009.
OLS、FWL 定理与遗漏变量偏误OLS, the Frisch–Waugh–Lovell theorem and omitted variable bias
OLS, the Frisch–Waugh–Lovell theorem, and omitted variable bias
「控制变量」到底在做什么?FWL 定理给出了一句话的答案:把它们的影响先剔掉。
What exactly does "controlling for X" do? FWL answers in one sentence: it strips X out first.
核心方程Core equations
展开逐步推导(2 步)Show the 2-step derivation
- 一阶条件First-order conditions$$\min_{\beta}(y-X\beta)'(y-X\beta)\ \Longrightarrow\ X'(y-X\hat\beta)=0$$正规方程的含义是「残差与每个回归元正交」,这是几何投影,与统计假设无关。The normal equations say that the residual is orthogonal to every regressor. That is geometry — a projection — and holds regardless of any statistical assumption.
- 无偏性需要外生性Unbiasedness requires exogeneity$$\hat\beta=\beta+(X'X)^{-1}X'u$$\(\mathbb{E}[u\mid X]=0\) 才能让第二项期望为零。同方差与无自相关只影响有效性与标准误,不影响无偏性——这两件事常被混为一谈。\(\mathbb{E}[u\mid X]=0\) is what makes the second term have zero expectation. Homoskedasticity and no autocorrelation bear on efficiency and standard errors, not on unbiasedness — the two are constantly conflated.
展开逐步推导(2 步)Show the 2-step derivation
- 三步做法Three steps$$\text{(1) }y\ \text{on}\ X_2\to\tilde y;\quad \text{(2) }x_1\ \text{on}\ X_2\to\tilde x_1;\quad \text{(3) }\tilde y\ \text{on}\ \tilde x_1$$得到的系数与三元回归里 \(x_1\) 的系数逐位相同(不是近似)。The coefficient obtained is numerically identical to the one from the full regression, not merely close to it.
- 这就是「控制」的定义This is what "controlling for" means$$\text{control }X_2\ \equiv\ \text{use only the part of }x_1\text{ orthogonal to }X_2$$固定效应就是控制个体虚拟变量的 FWL:等价于组内去均值。去趋势、季节调整同理,全是同一定理的应用。Fixed effects are FWL applied to a set of unit dummies: exactly equivalent to demeaning within groups. Detrending and seasonal adjustment are the same theorem again.
展开逐步推导(1 步)Show the 1-step derivation
- 符号判断Signing the bias$$\operatorname{sign}(\text{bias})=\operatorname{sign}(\beta_2)\times\operatorname{sign}(\delta)$$遗漏变量与被解释变量正相关、且与关键回归元正相关 ⇒ 高估。这是审稿人第一个会问的问题,也是最容易口头推断的诊断。If the omitted variable is positively related to the outcome and positively related to the regressor of interest, the estimate is too large. It is the first question a referee asks and the easiest diagnostic to run in your head.
交互图Interactive chart
教学算例Worked example
真值 \(y=1+2x_1-1.5x_2+\varepsilon\),其中 \(x_1=0.7x_2+\nu\)。
True model \(y=1+2x_1-1.5x_2+\varepsilon\), with \(x_1=0.7x_2+\nu\).
| 三元回归的 \(\hat\beta_1\)\(\hat\beta_1\) from the full regression | 2.0231 | |
| FWL 两步法FWL in two steps | 2.0231 | 逐位相同(差小于 1e−10)Identical to the digit (difference below 1e−10) |
| 漏掉 \(x_2\) 的估计Estimate omitting \(x_2\) | 1.2535 | |
| \(\hat\beta_1+\hat\beta_2\hat\delta\)\(\hat\beta_1+\hat\beta_2\hat\delta\) | 1.2535 | 遗漏变量偏误公式,恒等成立The omitted variable bias formula, an exact identity |
| 偏误大小Size of the bias | −0.7696 | \(\hat\beta_2\lt0\) 且 \(\hat\delta\gt0\) ⇒ 低估\(\hat\beta_2\lt0\) and \(\hat\delta\gt0\), so the estimate is too small |
经典文献Original sources
- Frisch, R. & Waugh, F. V. (1933). "Partial Time Regressions as Compared with Individual Trends." Econometrica 1(4), 387–401.原本是关于去趋势的争论:先去趋势再回归 vs 把趋势当回归元,二者等价。Originally an argument about detrending: removing a trend first and including it as a regressor turn out to be equivalent.
- Lovell, M. C. (1963). "Seasonal Adjustment of Economic Time Series." JASA 58(304), 993–1010.推广到季节调整,故称 FWL 定理。Extended to seasonal adjustment, whence the name FWL.
- Angrist, J. D. & Pischke, J.-S. (2009). Mostly Harmless Econometrics. Princeton University Press.把 FWL、遗漏变量偏误、IV、DID、RD 串成一条「因果识别」主线的现代标准入门。The modern standard introduction, running FWL, omitted variable bias, IV, DID and RD together as one argument about causal identification.
易错点Where it goes wrong
- 把「无偏」与「有效」搞混。高斯–马尔可夫的无偏性只要外生性;最小方差才需要同方差 + 无自相关。异方差不会让 OLS 有偏,只会让默认标准误错——所以对策是稳健标准误,而不是换估计量。Confusing unbiasedness with efficiency. The unbiasedness half of Gauss–Markov needs only exogeneity; minimum variance needs homoskedasticity and no autocorrelation. Heteroskedasticity does not bias OLS, it only invalidates the default standard errors — so the remedy is robust standard errors, not a different estimator.
- 只做一步残差化。FWL 要求 \(y\) 与 \(x_1\) 都对 \(X_2\) 残差化。只残差化一边,自由度与标准误都会错。Residualising only one side. FWL requires both \(y\) and \(x_1\) to be residualised against \(X_2\). Doing only one leaves the degrees of freedom and standard errors wrong.
- 用真值套遗漏变量偏误公式。有限样本里恒等式要用 \(\hat\beta_2\hat\delta\);\(\beta_2\delta\) 只在 plim 意义下成立。Plugging true values into the omitted variable bias formula. In a finite sample the identity requires \(\hat\beta_2\hat\delta\); \(\beta_2\delta\) holds only in probability limit.
工具变量与两阶段最小二乘Instrumental variables and two-stage least squares
Instrumental variables and 2SLS
找一个只通过 x 影响 y 的外生变异源,用它那一部分变异来识别因果。
Find a source of exogenous variation that reaches y only through x, and identify the causal effect from that part alone.
核心方程Core equations
展开逐步推导(3 步)Show the 3-step derivation
- 简单 IV 估计量The simple IV estimator$$\hat\beta_{IV}=\frac{\operatorname{Cov}(z,y)}{\operatorname{Cov}(z,x)}$$把 \(y=\beta x+u\) 代入:\(\dfrac{\beta\operatorname{Cov}(z,x)+\operatorname{Cov}(z,u)}{\operatorname{Cov}(z,x)}=\beta\),第二项为零全靠外生性。Substituting \(y=\beta x+u\) gives \(\dfrac{\beta\operatorname{Cov}(z,x)+\operatorname{Cov}(z,u)}{\operatorname{Cov}(z,x)}=\beta\), and the second term vanishes only by exogeneity.
- 两阶段最小二乘Two-stage least squares$$\text{(1) }x=\pi z+v\ \to\ \hat x;\qquad \text{(2) }y=\beta\hat x+e$$第一阶段把 \(x\) 投影到 \(z\) 的空间;第二阶段只用这块干净的变异。标准误必须用 2SLS 公式,手工分两步跑 OLS 会低估标准误。The first stage projects \(x\) onto the space spanned by \(z\); the second uses only that clean variation. The standard errors must come from the 2SLS formula — running two OLS regressions by hand understates them.
- 排他性约束不可检验The exclusion restriction cannot be tested$$\operatorname{Cov}(z,u)=0\ \text{is not testable when just-identified}$$这是 IV 最脆弱的地方:它是一个论证,不是一个检验。过度识别时的 Sargan/Hansen J 检验也只检验「若至少一个工具有效则其余有效」,不能自举。This is where IV is weakest: it is an argument, not a test. Even the Sargan/Hansen J test under over-identification only asks whether the instruments agree with one another, which cannot lift itself by its own bootstraps.
展开逐步推导(2 步)Show the 2-step derivation
- 弱工具的后果What weak instruments do$$\pi\to0\ \Longrightarrow\ \text{bias}\to\infty,\ \text{variance}\to\infty$$弱工具下 IV 比 OLS 更糟,且渐近分布不再正态。经验法则:第一阶段 F 统计量大于 10;但近年研究指出这个门槛远远不够。A weak instrument makes IV worse than OLS, and the asymptotic distribution ceases to be normal. The rule of thumb is a first-stage F above 10; recent work finds that threshold far too generous.
- 异质性效应下识别的是 LATEUnder heterogeneity, IV identifies the LATE$$\text{only for compliers}$$不是平均处理效应 ATE,也不是对全体人群的效应。换一个工具就换一批 complier,因此换一个 LATE——这解释了为什么不同 IV 研究的估计值差异很大。Not the average treatment effect, and not the effect for the population. A different instrument recruits a different set of compliers and therefore identifies a different LATE — which is why estimates across IV studies diverge so widely.
交互图Interactive chart
教学算例Worked example
\(y=x+u\),\(x=\pi z+0.8u+\text{噪声}\) ⇒ x 内生。
\(y=x+u\) with \(x=\pi z+0.8u+\text{noise}\), so x is endogenous.
| OLS(强工具情形)OLS (strong-instrument case) | 1.4881 | 偏高 49%——内生性偏误49% too high — the endogeneity bias |
| IV,π=0.8(强)IV, π=0.8 (strong) | 0.9984 | 接近真值Close to the truth |
| IV 标准差,π=0.8IV standard deviation, π=0.8 | 0.0410 | |
| IV,π=0.05(弱)IV, π=0.05 (weak) | −0.0888 | 完全崩坏,符号都反了Complete breakdown; even the sign is wrong |
| IV 标准差,π=0.05IV standard deviation, π=0.05 | 19.61 | 方差放大 478 倍Variance inflated 478-fold |
经典文献Original sources
- Wright, P. G. (1928). The Tariff on Animal and Vegetable Oils, Appendix B.IV 的最早出处。作者身份(是 Philip 还是其子 Sewall)至今仍有争议。The earliest known use of IV. Whether the author was Philip or his son Sewall is still disputed.
- Angrist, J. D. & Krueger, A. B. (1991). "Does Compulsory School Attendance Affect Schooling and Earnings?" QJE 106(4), 979–1014.用出生季度作教育年限的工具,1980 年人口普查约 32.9 万名 1930–39 年出生男性。经典中的经典,也是弱工具批评的经典靶子。Quarter of birth as an instrument for years of schooling, on roughly 329,000 men born 1930–39 in the 1980 US census. A classic, and equally a classic target for the weak-instrument critique.
- Bound, J., Jaeger, D. A. & Baker, R. M. (1995). "Problems with Instrumental Variables Estimation..." JASA 90(430), 443–450.用随机数当工具复现了 Angrist–Krueger 的结果——弱工具问题的著名演示。Reproduced the Angrist–Krueger results using randomly generated instruments — the celebrated demonstration of the weak-instrument problem.
- Imbens, G. W. & Angrist, J. D. (1994). "Identification and Estimation of Local Average Treatment Effects." Econometrica 62(2), 467–475.LATE 定理。Imbens 与 Angrist 因因果推断方法获 2021 年诺奖(与 Card 共享)。The LATE theorem. Imbens and Angrist shared the 2021 Nobel Prize with Card for work on causal inference.
易错点Where it goes wrong
- 手工跑两阶段 OLS 然后报第二阶段的标准误。那个标准误是错的(低估),因为它没考虑第一阶段的估计误差。必须用 2SLS/GMM 的联合公式。Running the two stages by hand and reporting the second-stage standard errors. Those standard errors are wrong (too small) because they ignore first-stage estimation error. Use the joint 2SLS or GMM formula.
- 把排他性约束当成可检验的假设。恰好识别时它完全不可检验,只能靠制度知识论证。「过度识别检验通过了」不等于工具有效。Treating the exclusion restriction as testable. Under just-identification it is not testable at all and can only be argued from institutional knowledge. Passing an over-identification test is not evidence that the instruments are valid.
- 把 LATE 当成 ATE 汇报。在效应异质时 IV 只识别 complier 的效应。政策外推前必须说清楚 complier 是谁。Reporting a LATE as though it were an ATE. Under heterogeneous effects IV identifies the effect for compliers only. Say who the compliers are before extrapolating to policy.
双重差分:从 2×2 到交错处理Difference-in-differences: from 2×2 to staggered adoption
Difference-in-differences: from 2×2 to staggered adoption
用「对照组的变化」当作「处理组本来会发生的变化」——以及这个做法近年是怎么被推翻重建的。
Use the change in the control group as the change the treated group would have seen — and see how that practice was overturned and rebuilt in recent years.
核心方程Core equations
展开逐步推导(2 步)Show the 2-step derivation
- 识别假设The identifying assumption$$\mathbb{E}[Y_{T,1}(0)-Y_{T,0}(0)]=\mathbb{E}[Y_{C,1}(0)-Y_{C,0}(0)]$$平行趋势针对的是反事实 Y(0),不是观测到的趋势。「处理前趋势平行」是支持性证据,不是假设本身——这个区分极其重要。Parallel trends is an assumption about the counterfactual Y(0), not about observed trends. Parallel pre-trends are supporting evidence, not the assumption itself — a distinction that matters enormously.
- 回归形式The regression form$$y_{it}=\alpha_i+\lambda_t+\tau D_{it}+\varepsilon_{it}$$个体固定效应吸收水平差异,时间固定效应吸收共同冲击。这就是模型 11 的 FWL 定理:\(\tau\) 用的是双向去均值后的残差变异。Unit fixed effects absorb level differences, time fixed effects absorb common shocks. This is the FWL theorem of Model 11: \(\tau\) is estimated from residual variation after two-way demeaning.
展开逐步推导(3 步)Show the 3-step derivation
- 分解The decomposition$$\text{TWFE}=\text{weighted average of all 2}\times\text{2 comparisons}$$其中包含一类「后处理组 vs 早处理组」的比较——把已经受处理的单位当成了对照组。Among those comparisons is a class that uses later-treated units against earlier-treated ones — that is, units that have already been treated serve as controls.
- 为什么会错Why this goes wrong$$\text{if effects grow over time, early-treated }y\text{ also rises}$$早处理组自己的增长被当成「共同趋势」减掉了。结果是系统性低估,权重甚至可能为负,导致符号翻转——即使每个个体的真实效应都为正。The growth in the early-treated group is subtracted out as though it were a common trend. The result is systematic understatement, with weights that can even be negative and flip the sign — while every individual treatment effect is positive.
- 现代做法What to do instead$$ATT(g,t):\ \text{group }g\text{ (first treated)}\times\text{time }t$$Callaway–Sant'Anna (2021) 只用干净的比较(对照组在两期都未受处理),因此在任意效应异质下仍然一致。事件研究图应当基于它,而不是简单的 TWFE 带 leads/lags。Callaway and Sant'Anna (2021) use only clean comparisons, in which the control group is untreated in both periods, and remain consistent under arbitrary heterogeneity. Event study plots should be built on this rather than on TWFE with leads and lags.
交互图Interactive chart
教学算例Worked example
1992 年新泽西最低工资由 4.25 美元升至 5.05 美元,宾夕法尼亚不变。快餐店全职等价(FTE)就业。
In 1992 New Jersey raised its minimum wage from $4.25 to $5.05 while Pennsylvania did not. Outcome: full-time-equivalent employment at fast-food restaurants.
| 新泽西 前 → 后New Jersey, before → after | 20.44 → 21.03 | 变化 +0.59Change +0.59 |
| 宾夕法尼亚 前 → 后Pennsylvania, before → after | 23.33 → 21.17 | 变化 −2.16Change −2.16 |
| DID 估计DID estimate | +2.75 | 就业不降反升,与竞争性劳动市场模型的预测相反Employment rose rather than fell, contrary to the competitive labour market prediction |
| —— 交错处理模拟 ———— staggered-timing simulation —— | ||
| 效应恒定时的 TWFETWFE with constant effects | 1.000 | 与真值一致,无偏Matches the truth; unbiased |
| 效应随时间增长时的 TWFETWFE with effects growing over time | 1.750 | |
| 同一情形的真实 ATTTrue ATT in the same setting | 2.417 | TWFE 低估 28%TWFE understates by 28% |
经典文献Original sources
- Card, D. & Krueger, A. B. (1994). "Minimum Wages and Employment: A Case Study of the Fast-Food Industry in New Jersey and Pennsylvania." American Economic Review 84(4), 772–793.Card 因劳动经济学的实证贡献获 2021 年诺奖,此文是引用核心。Card received the 2021 Nobel Prize for empirical work in labour economics, with this paper at the centre of the citation.
- Goodman-Bacon, A. (2021). "Difference-in-Differences with Variation in Treatment Timing." Journal of Econometrics 225(2), 254–277.把 TWFE 分解成所有 2×2 比较的加权平均,指出「后处理 vs 早处理」这一类比较的危害。Decomposes TWFE into a weighted average of all 2×2 comparisons and identifies the damage done by later-versus-earlier comparisons.
- Callaway, B. & Sant'Anna, P. H. C. (2021). "Difference-in-Differences with Multiple Time Periods." Journal of Econometrics 225(2), 200–230.group-time ATT(g,t) 估计量,只用干净比较,是目前的默认做法(R 包 did;Stata 命令 csdid)。The group-time ATT(g,t) estimator, built only from clean comparisons; now the default (R: did; Stata: csdid).
- de Chaisemartin, C. & D'Haultfœuille, X. (2020). "Two-Way Fixed Effects Estimators with Heterogeneous Treatment Effects." AER 110(9), 2964–2996.给出负权重的诊断方法,并提出替代估计量。Provides diagnostics for negative weights and an alternative estimator.
- Sun, L. & Abraham, S. (2021). "Estimating Dynamic Treatment Effects in Event Studies..." Journal of Econometrics 225(2), 175–199.事件研究图里 leads/lags 系数被污染的机制与修正。The mechanism by which leads and lags in event study plots are contaminated, and how to fix it.
易错点Where it goes wrong
- 把「处理前趋势平行」当成平行趋势假设成立的证明。假设针对的是处理后的反事实,本质不可检验。预趋势检验的功效往往很低,「看起来平行」的说服力被系统性高估。Treating parallel pre-trends as proof that parallel trends holds. The assumption concerns the post-treatment counterfactual and is fundamentally untestable. Pre-trend tests are often badly underpowered, and "it looks parallel" carries far less weight than it is given.
- 在交错处理下直接跑 TWFE 并报事件研究图。这是 2020 年以前的标准做法,现在已知有偏。应改用 Callaway–Sant'Anna 或 Sun–Abraham,并报告 Goodman-Bacon 分解看权重结构。Running TWFE under staggered timing and reporting the event study plot. Standard practice before 2020, now known to be biased. Use Callaway–Sant'Anna or Sun–Abraham instead, and report the Goodman-Bacon decomposition to inspect the weights.
- 对处理组数量很少的情形用常规聚类标准误。聚类数少于 40 左右时严重低估,应使用 wild cluster bootstrap 或随机化推断。Using conventional clustered standard errors with few treated clusters. Below roughly 40 clusters they understate badly; use a wild cluster bootstrap or randomisation inference.
断点回归设计Regression discontinuity design
Regression discontinuity design
在分数线两侧一分之差的人几乎一样——把这条线当作一次自然的随机分配。
People one mark either side of a cutoff are nearly identical — so treat the cutoff as a natural randomisation.
核心方程Core equations
展开逐步推导(3 步)Show the 3-step derivation
- 识别假设The identifying assumption$$\mathbb{E}[Y(0)\mid X=x],\ \mathbb{E}[Y(1)\mid X=x]\ \text{continuous at }x=c$$只要求潜在结果连续,不要求随机化。这是 RD 可信度高的根源:跳跃只能来自处理,因为别的一切都连续。Only continuity of potential outcomes is required, not randomisation. That is the source of the credibility: a jump can only come from treatment, because everything else is continuous.
- 局部随机化的解释The local randomisation reading$$\text{no precise manipulation of }X\ \Rightarrow\ \text{as-if random near }c$$Lee (2008) 的论证。对应的检验是 McCrary 密度检验:看 \(X\) 的密度在断点处是否有跳跃,有跳跃说明有人在操纵。Lee's (2008) argument. The corresponding check is the McCrary density test: a jump in the density of \(X\) at the cutoff indicates that someone is manipulating it.
- 模糊 RDFuzzy RD$$\tau_{FRD}=\frac{\text{jump in outcome}}{\text{jump in treatment probability}}$$这就是以「过线与否」为工具的 IV——分子分母都是跳跃,形式与 Wald 估计量一致。This is IV with "crossing the threshold" as the instrument — numerator and denominator are both jumps, in the form of a Wald estimator.
展开逐步推导(3 步)Show the 3-step derivation
- 窄带宽Narrow bandwidth$$h\downarrow:\ \text{bias}\downarrow,\ \text{variance}\uparrow$$样本少了,噪声大。Fewer observations, hence more noise.
- 宽带宽Wide bandwidth$$h\uparrow:\ \text{bias}\uparrow,\ \text{variance}\downarrow$$把远处的点也拿来外推,函数弯曲就会被当成跳跃。Distant points are brought in and extrapolated, so curvature in the function is read as a jump.
- 一个反直觉的细节A counter-intuitive detail$$\text{equal curvature on both sides}\ \Rightarrow\ \text{extrapolation bias cancels}$$所以偏误来自「两侧曲率不同」,不是来自「有曲率」。这一点是写验算脚本时发现的:用对称的 \(0.8x^2\) 做模拟,宽带宽根本看不出偏误。So the bias arises from curvature that differs across the two sides, not from curvature as such. This surfaced while writing the verification script: with a symmetric \(0.8x^2\), a wide bandwidth shows no bias at all.
交互图Interactive chart
教学算例Worked example
\(x\sim U[-1,1]\),\(n=4000\),噪声 \(\sigma=0.5\);两侧曲率不同(左 0.4,右 2.5)。
\(x\sim U[-1,1]\), \(n=4000\), noise \(\sigma=0.5\); curvature differs across sides (0.4 left, 2.5 right).
| h = 0.10h = 0.10 | 0.804 (se 0.104, n=381) | 方差大High variance |
| h = 0.30h = 0.30 | 0.919 (se 0.060, n=1149) | 偏误与方差的较好折中The better compromise between bias and variance |
| h = 0.90h = 0.90 | 0.689 (se 0.035, n=3587) | 标准误最小,但偏误最大Smallest standard error, largest bias |
| 若两侧曲率相同With equal curvature on both sides | 宽带宽也几乎无偏 | 外推偏误相互抵消Even a wide bandwidth is nearly unbiased; the extrapolation biases cancel |
经典文献Original sources
- Thistlethwaite, D. L. & Campbell, D. T. (1960). "Regression-Discontinuity Analysis: An Alternative to the Ex Post Facto Experiment." Journal of Educational Psychology 51(6), 309–317.RD 的原始文献,研究国家优秀学生奖学金对后续学业的影响。此后沉寂了近 40 年。The original RD paper, on the effect of National Merit awards on later academic outcomes. It then lay dormant for close to forty years.
- Hahn, J., Todd, P. & van der Klaauw, W. (2001). "Identification and Estimation of Treatment Effects with a Regression-Discontinuity Design." Econometrica 69(1), 201–209.现代 RD 的理论基础:给出连续性识别条件与局部线性估计。The theoretical basis of modern RD: continuity-based identification and local linear estimation.
- Lee, D. S. (2008). "Randomized Experiments from Non-random Selection in U.S. House Elections." Journal of Econometrics 142(2), 675–697.局部随机化解释;美国众议院现任优势的经典 RD 应用。The local randomisation reading, and the classic RD application to incumbency advantage in US House elections.
- Calonico, S., Cattaneo, M. D. & Titiunik, R. (2014). "Robust Nonparametric Confidence Intervals for Regression-Discontinuity Designs." Econometrica 82(6), 2295–2326.偏误修正的稳健置信区间,现在的标准做法(rdrobust)。Bias-corrected robust confidence intervals, now standard practice (rdrobust).
- McCrary, J. (2008). "Manipulation of the Running Variable in the Regression Discontinuity Design." Journal of Econometrics 142(2), 698–714.密度检验,判断有没有人在操纵分数。The density test for manipulation of the running variable.
易错点Where it goes wrong
- 用高阶全局多项式拟合。Gelman–Imbens (2019) 证明三次以上的全局多项式会给出荒谬的权重,导致伪跳跃。现在的标准是局部线性 + MSE 最优带宽 + 偏误修正置信区间。Fitting high-order global polynomials. Gelman and Imbens (2019) show that global polynomials of order three and above produce absurd implicit weights and spurious jumps. The standard is now local linear estimation with an MSE-optimal bandwidth and bias-corrected confidence intervals.
- 不做 McCrary 密度检验与协变量平衡检验。如果个体能精确操纵 running variable,断点两侧就不再可比,RD 的全部可信度立刻消失。Omitting the McCrary density test and covariate balance checks. If units can manipulate the running variable precisely, the two sides of the cutoff are no longer comparable and the credibility of RD evaporates at once.
- 把 RD 估计外推到全体人群。它识别的只是断点处的处理效应。分数线附近的学生与远离分数线的学生,效应可能完全不同。Extrapolating the RD estimate to the population. What is identified is the treatment effect at the cutoff. Students near the threshold and students far from it may respond quite differently.
GMM 与时间序列:平稳性、单位根、协整GMM and time series: stationarity, unit roots, cointegration
GMM, stationarity, unit roots, and cointegration
所有估计量都是矩条件的解;而时间序列的第一个问题永远是:这个序列平稳吗?
Every estimator is the solution to a set of moment conditions; and the first question in time series is always whether the series is stationary.
核心方程Core equations
展开逐步推导(3 步)Show the 3-step derivation
- 常见估计量都是特例Familiar estimators as special cases$$\text{OLS: }\mathbb{E}[x_iu_i]=0;\quad \text{IV: }\mathbb{E}[z_iu_i]=0;\quad \text{MLE: }\mathbb{E}[\partial\ln f/\partial\theta]=0$$换矩条件就换估计量,这是 GMM 的全部威力。Change the moment conditions and you change the estimator. That is the whole power of GMM.
- 最优权重矩阵The optimal weighting matrix$$W^*=\bigl(\mathbb{E}[gg']\bigr)^{-1}=\Omega^{-1}$$给方差小的矩条件更大权重。两步 GMM:先用 \(W=I\) 得一致估计,再估 \(\Omega\),再重估。Moment conditions with smaller variance receive more weight. Two-step GMM: obtain a consistent estimate with \(W=I\), estimate \(\Omega\), re-estimate.
- 过度识别检验The over-identification test$$J=n\,\bar g(\hat\theta)'\hat\Omega^{-1}\bar g(\hat\theta)\ \xrightarrow{d}\ \chi^2_{m-k}$$\(m\) 个矩条件、\(k\) 个参数 ⇒ 自由度 \(m-k\)。恰好识别时 J 恒等于 0,检验不存在——这就是为什么排他性约束在恰好识别的 IV 里不可检验。\(m\) moment conditions and \(k\) parameters give \(m-k\) degrees of freedom. Under just-identification J is identically zero and the test does not exist — which is exactly why the exclusion restriction is untestable in a just-identified IV.
展开逐步推导(3 步)Show the 3-step derivation
- 平稳时的方差与半衰期Variance and half-life under stationarity$$\operatorname{Var}(y)=\frac{\sigma^2}{1-\rho^2},\qquad \text{half-life}=\frac{\ln 0.5}{\ln\rho}$$\(\rho\to1\) 时方差发散,冲击永不消退。As \(\rho\to1\) the variance diverges and shocks never die away.
- 为什么不能用常规 t 检验Why the usual t test fails$$\text{under }\rho=1,\ \hat\rho\ \text{has a non-normal (Wiener) limit}$$所以要用 Dickey–Fuller 的专用临界值,比常规 t 临界值更负。用常规临界值会过度拒绝单位根。Hence the dedicated Dickey–Fuller critical values, which lie further into the negative tail. Using conventional critical values over-rejects the unit root.
- 协整Cointegration$$y_t,x_t\sim I(1)\ \text{but}\ y_t-\beta x_t\sim I(0)$$两个各自游走的序列被一条长期关系拴住。Granger 表示定理:协整 ⟺ 存在误差修正模型 \(\Delta y_t=\alpha(y_{t-1}-\beta x_{t-1})+\dots\)。Two series that each wander are tied together by a long-run relation. Granger's representation theorem: cointegration holds if and only if an error correction model exists, \(\Delta y_t=\alpha(y_{t-1}-\beta x_{t-1})+\dots\).
交互图Interactive chart
教学算例Worked example
\(y_t=\rho y_{t-1}+\varepsilon_t\),\(\varepsilon\sim N(0,1)\),固定随机种子。
\(y_t=\rho y_{t-1}+\varepsilon_t\), \(\varepsilon\sim N(0,1)\), fixed random seed.
| ρ = 0.5ρ = 0.5 | 样本方差 1.12 | 理论值 \(1/(1-0.25)=1.33\)Theoretical value \(1/(1-0.25)=1.33\) |
| ρ = 0.9ρ = 0.9 | 样本方差 3.47 | 理论值 \(1/(1-0.81)=5.26\)Theoretical value \(1/(1-0.81)=5.26\) |
| ρ = 1.0ρ = 1.0 | 样本方差 15.01 | 理论方差不存在——随 T 增长而发散No theoretical variance exists — it diverges with T |
| ρ=0.9 的冲击半衰期Half-life of a shock at ρ=0.9 | 6.58 期 | \(\ln0.5/\ln0.9\)\(\ln0.5/\ln0.9\) |
| 3 个矩条件、1 个参数3 moment conditions, 1 parameter | J 检验自由度 = 2 | J test with 2 degrees of freedom |
经典文献Original sources
- Hansen, L. P. (1982). "Large Sample Properties of Generalized Method of Moments Estimators." Econometrica 50(4), 1029–1054.GMM 的奠基文献。2013 年诺奖。The founding paper of GMM. Nobel Prize 2013.
- Dickey, D. A. & Fuller, W. A. (1979). "Distribution of the Estimators for Autoregressive Time Series with a Unit Root." JASA 74(366), 427–431.单位根检验与非标准渐近分布。Unit root tests and their non-standard asymptotic distribution.
- Granger, C. W. J. & Newbold, P. (1974). "Spurious Regressions in Econometrics." Journal of Econometrics 2(2), 111–120.两个独立随机游走互相回归,R² 却很高——伪回归问题的著名演示。Two independent random walks regressed on each other, with a high R² — the celebrated demonstration of spurious regression.
- Engle, R. F. & Granger, C. W. J. (1987). "Co-integration and Error Correction." Econometrica 55(2), 251–276.协整与误差修正模型。二人共获 2003 年诺奖(Engle 另因 ARCH)。Cointegration and error correction. The two shared the 2003 Nobel Prize (Engle also for ARCH).
易错点Where it goes wrong
- 对不平稳序列直接跑回归。两个独立的随机游走互相回归会得到很高的 R² 与很大的 t 值,全是假的。先做单位根检验,再决定是差分还是建协整/误差修正模型。Regressing non-stationary series directly. Two independent random walks regressed on one another yield a high R² and large t-statistics, all of them false. Test for unit roots first, then decide between differencing and a cointegration or error correction model.
- 差分掉一切以求平稳。若序列本来协整,差分会丢掉长期关系的信息,得到只有短期动态的模型。协整时正确做法是误差修正模型,不是无脑差分。Differencing everything in pursuit of stationarity. If the series are cointegrated, differencing discards the long-run relation and leaves a model of short-run dynamics alone. Under cointegration the correct response is an error correction model, not differencing on reflex.
- 在恰好识别时报告 J 检验。自由度为 0,\(J\) 恒等于 0,这个「检验」没有任何信息量。Reporting a J test under just-identification. With zero degrees of freedom \(J\) is identically zero and the test carries no information whatever.