統計で英語多読 4-4: 正規分布 — 統計学で最も重要な「ベルカーブ」
正規分布を題材にした英語多読ユニット。約540語の英文と全文日本語訳で、ベルカーブの性質と68-95-99.7ルールを読みます。
「大人のための英語多読図書館」へようこそ。今回は、統計学の主役とも言える正規分布(ベルカーブ)の形を決める2つのパラメータと、便利な「68-95-99.7ルール」を、英語の文章でたどっていきます。
📊 このユニットの情報 語数: 約540語 / 推定読了時間: 5〜6分 / 難易度: ★★☆☆☆(初中級)
Learning Objectives
みなさんこんにちは。大人のための英語多読図書館へようこそ。
今回は、ついに統計学の主役とも言える「正規分布」の登場です。英語では Normal Distribution と言いますが、その美しい左右対称の形から「ベルカーブ(bell curve)」という愛称で親しまれています。
なぜ正規分布がこれほど重要なのでしょうか? それは、身長や体重、テストのスコア、測定誤差など、自然界や社会における非常に多くのデータが、この正規分布によく従うからです。
この記事では、正規分布の形が何によって決まるのか、そしてデータが分布の中でどのあたりに位置するのかを大まかに把握できる、とても便利な「68-95-99.7ルール」について学びます。
それでは今回も多読を楽しんでいきましょう。
Summary
- The Bell Curve: The normal distribution is a symmetric, bell-shaped probability distribution that is fundamental in statistics.
- Ubiquity: Many natural and social phenomena, such as height, blood pressure, and measurement errors, can be approximated by the normal distribution.
- Parameters: The shape and position of the normal distribution are determined by two parameters: the mean (\(\mu\)) and the standard deviation (\(\sigma\)).
- The 68-95-99.7 Rule: This is an empirical rule that describes where the data lies in a normal distribution:
- About 68% of data falls within 1 standard deviation of the mean.
- About 95% of data falls within 2 standard deviations of the mean.
- About 99.7% of data falls within 3 standard deviations of the mean.
Explanation
The Most Important Distribution in Statistics
The normal distribution, also known as the Gaussian distribution, is arguably the most important concept in all of statistics. Why? Because it appears so frequently in the real world. Many continuous variables, when you collect enough data, tend to follow this bell-shaped pattern.
The curve of a normal distribution has several key characteristics: - It is bell-shaped and perfectly symmetric around its center. - The mean, median, and mode are all equal and located at the center of the distribution. - The curve extends indefinitely in both directions, getting closer and closer to the horizontal axis but never touching it.
The Two Parameters: Mean and Standard Deviation
The exact shape of the bell curve is defined by two parameters: 1. The Mean (\(\mu\)): This determines the location of the center of the distribution. If you change the mean, the entire curve shifts to the left or right along the x-axis. 2. The Standard Deviation (\(\sigma\)): This determines the spread or width of the distribution. - A small standard deviation results in a tall, narrow curve, meaning most data points are clustered closely around the mean. - A large standard deviation results in a short, wide curve, meaning the data is more spread out.
While the mathematical formula for the normal distribution's PDF is quite complex, you don't need to memorize it. The crucial takeaway is that \(\mu\) and \(\sigma\) are the only two numbers you need to define any normal distribution.
A Powerful Rule of Thumb: The 68-95-99.7 Rule
One of the most practical features of the normal distribution is the Empirical Rule, or the 68-95-99.7 rule. This rule provides a quick way to understand the spread of data without complex calculations.
It states that for any normal distribution, regardless of its mean and standard deviation: - Approximately 68% of the data points lie within one standard deviation of the mean (i.e., in the range from \(\mu - \sigma\) to \(\mu + \sigma\)). - Approximately 95% of the data points lie within two standard deviations of the mean (from \(\mu - 2\sigma\) to \(\mu + 2\sigma\)). - Approximately 99.7% of the data points lie within three standard deviations of the mean (from \(\mu - 3\sigma\) to \(\mu + 3\sigma\)).
Let's say the IQ scores of a population are normally distributed with a mean of 100 and a standard deviation of 15. Using this rule, we can quickly say that: - About 68% of people have an IQ between 85 (\(100-15\)) and 115 (\(100+15\)). - About 95% of people have an IQ between 70 (\(100-30\)) and 130 (\(100+30\)). - Almost everyone (99.7%) has an IQ between 55 and 145.
This rule is incredibly useful for getting a feel for your data and identifying values that are common versus those that are unusual or outliers.
まとめ
今回は、統計学で最も重要な「正規分布(ベルカーブ)」を取り上げました。左右対称の釣り鐘型で、平均・中央値・最頻値がすべて中心で一致するこの分布は、その形が平均(\(\mu\))と標準偏差(\(\sigma\))という2つのパラメータだけで決まります。さらに、データの約68%・95%・99.7%が平均からそれぞれ1・2・3標準偏差の範囲に収まるという「68-95-99.7ルール」を使えば、複雑な計算なしにデータの散らばり具合を素早くつかめることを見てきました。
次回は、あらゆる正規分布を平均0・標準偏差1の「標準正規分布」に変換する標準化と、ZスコアやZ表を使った確率計算の方法を見ていきます。
日本語訳(全文)
英文を最後まで読み終えてから、答え合わせ用にお使いください。多読の原則として、まずは訳を見ずに英文だけで理解を試みることをおすすめします。
Summary
- ベルカーブ: 正規分布は、左右対称の釣り鐘型をした確率分布で、統計学において基礎的な存在です。
- 遍在性: 身長、血圧、測定誤差など、自然界や社会の多くの現象は、正規分布で近似することができます。
- パラメータ: 正規分布の形と位置は、平均(\(\mu\))と標準偏差(\(\sigma\))という2つのパラメータによって決まります。
- 68-95-99.7ルール: これは、正規分布においてデータがどこに位置するかを表す経験則です。
- データの約68%は、平均から1標準偏差の範囲内に収まります。
- データの約95%は、平均から2標準偏差の範囲内に収まります。
- データの約99.7%は、平均から3標準偏差の範囲内に収まります。
統計学で最も重要な分布
正規分布は、ガウス分布としても知られ、統計学全体のなかでおそらく最も重要な概念です。なぜでしょうか。それは、現実の世界に非常に頻繁に現れるからです。多くの連続的な変数は、十分なデータを集めると、この釣り鐘型のパターンに従う傾向があります。
正規分布の曲線には、いくつかの重要な特徴があります。 - 釣り鐘型で、その中心を軸として完全に左右対称です。 - 平均、中央値、最頻値がすべて等しく、分布の中心に位置します。 - 曲線は両方向に限りなく伸びていき、横軸に近づき続けますが、決して触れることはありません。
2つのパラメータ: 平均と標準偏差
ベルカーブの正確な形は、2つのパラメータによって定義されます。 1. 平均(\(\mu\)): これは分布の中心の位置を決めます。平均を変えると、曲線全体がx軸に沿って左または右へ移動します。 2. 標準偏差(\(\sigma\)): これは分布の散らばり、つまり幅を決めます。 - 標準偏差が小さいと、背が高く幅の狭い曲線になります。これは、ほとんどのデータ点が平均の近くに密集していることを意味します。 - 標準偏差が大きいと、背が低く幅の広い曲線になります。これは、データがより広く散らばっていることを意味します。
正規分布の確率密度関数(PDF)の数式はかなり複雑ですが、それを暗記する必要はありません。重要な要点は、どんな正規分布を定義するのにも必要な数値は \(\mu\) と \(\sigma\) の2つだけだ、ということです。
強力な経験則: 68-95-99.7ルール
正規分布の最も実用的な特徴の一つが、経験則、すなわち68-95-99.7ルールです。このルールは、複雑な計算なしにデータの散らばりを理解する素早い方法を与えてくれます。
このルールは、どんな正規分布についても、その平均と標準偏差にかかわらず、次のことが成り立つと述べています。 - データ点の約68%は、平均から1標準偏差の範囲内に収まります(すなわち、\(\mu - \sigma\) から \(\mu + \sigma\) までの範囲)。 - データ点の約95%は、平均から2標準偏差の範囲内に収まります(\(\mu - 2\sigma\) から \(\mu + 2\sigma\) まで)。 - データ点の約99.7%は、平均から3標準偏差の範囲内に収まります(\(\mu - 3\sigma\) から \(\mu + 3\sigma\) まで)。
ある母集団のIQスコアが、平均100、標準偏差15の正規分布に従うとしましょう。このルールを使えば、次のように素早く言うことができます。 - 約68%の人が、85(\(100-15\))から115(\(100+15\))の間のIQを持っています。 - 約95%の人が、70(\(100-30\))から130(\(100+30\))の間のIQを持っています。 - ほとんどすべての人(99.7%)が、55から145の間のIQを持っています。
このルールは、データの感覚をつかんだり、ありふれた値とそうでない珍しい値や外れ値を見分けたりするのに、とても役立ちます。