統計で英語多読 4-6: 大数の法則と中心極限定理 — 推測統計学の2大支柱
大数の法則と中心極限定理を題材にした英語多読ユニット。約620語の英文と全文日本語訳で、推測統計を支える2大定理の意味を読みます。
「大人のための英語多読図書館」へようこそ。今回は、標本から母集団を推測する推測統計学を支える「大数の法則」と「中心極限定理」という2つの定理を、英語の文章でたどっていきます。
📊 このユニットの情報 語数: 約620語 / 推定読了時間: 5〜7分 / 難易度: ★★☆☆☆(初中級)
Learning Objectives
みなさんこんにちは。大人のための英語多読図書館へようこそ。
今回は、一見すると難しそうですが、標本調査から社会全体を推測する「推測統計学」の根幹を支える、2つの非常に重要な定理を学びます。
-
大数の法則 (Law of Large Numbers): これは、「試行回数を増やせば増やすほど、その平均値は『真の平均値』に近づいていく」という法則です。例えば、サイコロを数回振っただけでは出る目の平均はバラバラですが、何万回も振れば、その平均は期待値である3.5に限りなく近づきます。保険会社がビジネスを成り立たせることができるのも、この法則のおかげです。
-
中心極限定理 (Central Limit Theorem): こちらはさらに強力で、少し不思議な定理です。「元のデータがどんな形の分布であっても、そこからたくさんのサンプルを取ってきて、その『サンプルたちの平均値』の分布を調べると、なぜか正規分布に近づく」というものです。この定理があるからこそ、私たちは世の中の多くの現象を正規分布を使って分析できるのです。
この2つの定理は、私たちがサンプルデータから母集団全体について自信を持って語ることを可能にする、理論的な「お守り」のようなものです。
それでは今回も多読を楽しんでいきましょう。
Summary
- Inferential Statistics: These two theorems are the theoretical foundation that allows us to make inferences about a population from a sample.
- Law of Large Numbers (LLN): As the sample size (
n) gets larger, the sample mean (\(\bar{x}\)) gets closer to the true population mean (\(\mu\)). It guarantees that sampling is a reliable way to estimate the population average in the long run. - Central Limit Theorem (CLT): Regardless of the population's original distribution, the distribution of the sample means will be approximately normal, as long as the sample size is sufficiently large (usually n > 30).
- Why CLT is Powerful: The CLT allows us to use normal distribution theory for statistical inference (like constructing confidence intervals and hypothesis testing) even when we don't know the shape of the population distribution.
Explanation
The Foundation of Inferential Statistics
So far, we have learned about different probability distributions. Now, we will discuss two of the most important theorems in all of statistics: the Law of Large Numbers and the Central Limit Theorem. These theorems form the bedrock of inferential statistics, which is the science of drawing conclusions about a whole population based on a smaller sample.
The Law of Large Numbers (LLN): Getting Closer to the Truth
The Law of Large Numbers is an intuitive but powerful idea. It states that as you increase the number of trials or observations in an experiment, the average result from those trials will get closer and closer to the expected value (the true mean of the population).
Let's use the example of rolling a fair six-sided die. The true mean (expected value) of a single roll is 3.5. If you roll it only 5 times, your average might be something like 2.8 or 4.6—not very close to 3.5. However, the LLN tells us that if you roll that same die 10,000 times, the average of all those rolls will be extremely close to 3.5.
This principle is what makes casinos and insurance companies profitable. They may lose money on a single event, but over millions of events, the average outcome becomes highly predictable, allowing them to set prices that guarantee a long-term profit. The LLN gives us confidence that a large sample will provide a good estimate of the population mean.
The Central Limit Theorem (CLT): The Magic of the Normal Distribution
The Central Limit Theorem (CLT) is even more remarkable and perhaps less intuitive. It is the reason why the normal distribution is so central to statistics.
The CLT states the following: - Take any population, no matter what its distribution looks like (it could be uniform, skewed, or anything else). - Repeatedly draw random samples of a sufficiently large size (a common rule of thumb is \(n > 30\)) from this population. - Calculate the mean of each of these samples. - The distribution of these sample means will be approximately a normal distribution.
This is an astonishing result. Even if you start with a population that looks nothing like a bell curve, the distribution of its sample means will magically form a bell curve.
Furthermore, the mean of this distribution of sample means will be equal to the original population mean (\(\mu\)). And the standard deviation of this distribution of sample means (called the "standard error") will be the population standard deviation divided by the square root of the sample size (\(\sigma / \sqrt{n}\)).
Why is this so important? The CLT allows us to apply the methods of normal distribution theory—like calculating probabilities with Z-scores, creating confidence intervals, and performing hypothesis tests—to the sample mean, even if we have no idea what the original population's distribution looks like. It is the bridge that connects the data we have (the sample) to the population we want to understand, making much of modern statistical analysis possible.
まとめ
今回は、標本から母集団を推測する推測統計学を理論的に支える2つの定理を学びました。「大数の法則」は、試行回数を増やすほど標本平均が真の母平均に近づくという法則で、サンプリングが母平均を推定する信頼できる方法であることを保証してくれます。一方の「中心極限定理」は、元の母集団がどんな分布であっても、十分大きな標本を取れば標本平均の分布が正規分布に近づくという、より強力で不思議な定理です。この2つが、手元の標本と知りたい母集団とを結ぶ橋渡しとなり、現代の統計分析を可能にしていることを見てきました。
次回は、良い推測を行うための出発点として、母集団から標本をどう選ぶかという「標本抽出法(サンプリング)」を見ていきます。
日本語訳(全文)
英文を最後まで読み終えてから、答え合わせ用にお使いください。多読の原則として、まずは訳を見ずに英文だけで理解を試みることをおすすめします。
Summary
- 推測統計学: これら2つの定理は、標本から母集団について推論を行うことを可能にする理論的な基盤です。
- 大数の法則(LLN): 標本サイズ(
n)が大きくなるにつれて、標本平均(\(\bar{x}\))は真の母平均(\(\mu\))に近づきます。これは、長期的に見れば、標本抽出が母集団の平均を推定する信頼できる方法であることを保証してくれます。 - 中心極限定理(CLT): 母集団の元の分布がどうであれ、標本サイズが十分に大きければ(通常は n > 30)、標本平均の分布はほぼ正規分布になります。
- CLTが強力な理由: CLTのおかげで、母集団の分布の形を知らなくても、統計的推論(信頼区間の構築や仮説検定など)に正規分布の理論を使うことができます。
推測統計学の基盤
ここまで、私たちはさまざまな確率分布について学んできました。これから、統計学全体のなかで最も重要な2つの定理、すなわち大数の法則と中心極限定理について論じます。これらの定理は、より小さな標本にもとづいて母集団全体についての結論を導く科学である、推測統計学の土台をなしています。
大数の法則(LLN): 真実に近づいていく
大数の法則は、直感的でありながら強力な考え方です。これは、実験における試行や観測の回数を増やすほど、それらの試行から得られる平均値が、期待値(母集団の真の平均)にどんどん近づいていく、というものです。
公平な6面のサイコロを振る例を使ってみましょう。1回振ったときの真の平均(期待値)は3.5です。もし5回だけ振ったなら、その平均は2.8や4.6のような、3.5にあまり近くない値になるかもしれません。しかし大数の法則は、同じサイコロを10,000回振れば、それらすべての出目の平均が3.5に極めて近くなる、と教えてくれます。
この原理こそが、カジノや保険会社を利益の出るものにしているのです。彼らは1回の出来事では損をするかもしれませんが、何百万回もの出来事を通じてみれば、平均的な結果は非常に予測しやすくなり、長期的な利益を保証する価格を設定できるようになります。大数の法則は、大きな標本が母平均の良い推定値を与えてくれるという確信を、私たちに与えてくれます。
中心極限定理(CLT): 正規分布の魔法
中心極限定理(CLT)は、さらに注目に値する、そしておそらくより直感に反する定理です。これこそが、正規分布が統計学においてこれほど中心的である理由です。
CLTは次のように述べています。 - その分布がどんな形をしていようとも(一様でも、ゆがんでいても、その他なんでも)、どんな母集団でもよいので、それを取り上げます。 - その母集団から、十分に大きいサイズ(よく使われる目安は \(n > 30\))のランダムな標本を、繰り返し抽出します。 - これらの標本それぞれの平均を計算します。 - これらの標本平均の分布は、ほぼ正規分布になります。
これは驚くべき結果です。たとえベルカーブにまったく似ていない母集団から始めたとしても、その標本平均の分布は、魔法のようにベルカーブを形づくるのです。
さらに、この標本平均の分布の平均は、元の母平均(\(\mu\))に等しくなります。そして、この標本平均の分布の標準偏差(「標準誤差」と呼ばれます)は、母集団の標準偏差を標本サイズの平方根で割ったもの(\(\sigma / \sqrt{n}\))になります。
なぜこれがそれほど重要なのでしょうか。CLTのおかげで、元の母集団の分布がどんな形なのかまったく分からなくても、Zスコアによる確率計算、信頼区間の構築、仮説検定の実施といった正規分布の理論の手法を、標本平均に対して適用できるのです。それは、手元にあるデータ(標本)と、理解したい母集団とを結ぶ橋であり、現代の統計分析の多くを可能にしているのです。