英語多読

統計で英語多読 1-2: 母集団と標本 — 推測の第一歩

母集団と標本を題材にした英語多読ユニット。約840語の英文と全文日本語訳で、標本調査の利点、無作為抽出と標本誤差の考え方をたどります。

大人のための英語多読図書館」へようこそ。今回は、調べたい対象全体である「母集団」と、実際に調査する一部である「標本」という統計の出発点を、無作為抽出や標本誤差の考え方とあわせて、英語の文章でたどっていきます。

📊 このユニットの情報 語数: 約840語 / 推定読了時間: 6〜9分 / 難易度: ★★☆☆☆(初中級)

Learning Objectives

みなさんこんにちは。大人のための英語多読図書館へようこそ。

今回は、統計学の基本的な考え方である「母集団」と「標本」について学びます。この2つの言葉は、統計的な推測を行う上での出発点となる、とても重要なコンセプトです。

例えば、日本の有権者全員の「内閣支持率」を知りたいとします。でも、1億人以上の有権者全員に意見を聞いて回るのは、時間もお金もかかりすぎて現実的ではありませんよね。そこで、私たちは何をするでしょうか?そう、ニュースでよく見るように、1000人や2000人といった一部の人に電話調査などをして、その結果から全体の支持率を「推測」します。

このとき、調査の対象となる「日本の有権者全体」のことを母集団と呼び、実際に調査した「1000人の有権者」のことを標本(サンプル)と呼びます。

この記事では、なぜ標本調査が必要なのか、そして、どうすれば標本から母集団のことを正しく推測できるのか、その鍵となる「無作為抽出」の重要性や、推測に伴う「標本誤差」という考え方について解説していきます。

それでは今回も多読を楽しんでいきましょう。

Summary

  • Population vs. Sample: The population is the entire group you want to study or draw conclusions about (e.g., all voters in a country). A sample is a smaller, manageable subset of that population from which you actually collect data (e.g., 1,000 voters selected for a poll).
  • Census vs. Sample Survey: A census collects data from every single member of the population. A sample survey collects data from only a sample. We use sample surveys because they are cheaper, faster, and often more practical than a census.
  • Importance of Random Sampling: To make accurate inferences about the population, the sample must be representative. Random sampling, where every member of the population has an equal chance of being selected, is the best way to avoid bias and achieve a representative sample.
  • Sampling Error: Sampling error is the natural difference between a statistic calculated from a sample (e.g., the average opinion of the sample) and the true parameter of the population (the average opinion of everyone). It's not a "mistake," but an unavoidable consequence of not studying the entire population. Larger, well-chosen samples tend to have smaller sampling errors.

Explanation

The Big Picture: Population and Sample

In statistics, we often want to understand something about a large group of individuals or items. This entire group is called the population. A population can be very broad, like "all high school students in Japan," or more specific, like "all light bulbs produced at a certain factory."

However, studying every single member of a population is often difficult or impossible. This is called a census. While a country's national census tries to do this, it's a massive undertaking that takes years and costs a huge amount of money. For most research and business questions, a census is simply not practical.

So, what's the solution? We select a smaller group from the population to study. This smaller group is called a sample. We then collect data from this sample and use the findings to make conclusions about the entire population. This process is called a sample survey. For example, a TV ratings company can't monitor what every single household is watching. Instead, they select a sample of a few thousand households and use their viewing habits to estimate the ratings for the entire country.

Why Do We Use Samples?

There are several key advantages to using a sample survey over a census:

  • Cost-Effective: It's much cheaper to collect data from a few thousand people than from millions.
  • Time-Saving: Analyzing a smaller dataset is significantly faster. This allows us to get timely results, which is crucial for things like political polls or economic reports.
  • Practicality: Sometimes, the process of collecting data destroys the item being tested. For example, if a company wants to test the lifespan of its batteries, they can't test every single one until it runs out. They must use a sample.

The Key to Good Inference: Random Sampling

The goal of using a sample is to make an accurate inference, or guess, about the population. For this to work, the sample must be representative of the population. This means the characteristics of the people in the sample should closely match the characteristics of the people in the population.

How do we get a representative sample? The best method is random sampling. In a simple random sample, every member of the population has an equal chance of being selected. This is like putting everyone's name into a giant hat and drawing names out.

Random sampling is crucial because it helps to avoid bias. Bias is a systematic error that makes our sample unrepresentative. For example, if you wanted to know the average height of students at a university and only surveyed the basketball team, your sample would be biased, and your results would be inaccurate. Random sampling ensures that tall and short students alike have a chance to be included, giving us a much more accurate picture of the whole university.

Understanding Sampling Error

Even with perfect random sampling, the results from our sample will almost never be exactly the same as the results from the entire population. Imagine you want to know the average age of people in a town. The true average age of all residents is the population parameter. Now, you take a random sample of 100 people and calculate their average age. This is the sample statistic.

If you took another random sample of 100 people, you would likely get a slightly different average age. This natural variation from sample to sample is called sampling error. It's important to understand that sampling error is not a "mistake" in your methods. It's the unavoidable discrepancy that arises because a sample is just a part of the whole population.

We can reduce sampling error by increasing the sample size. A larger sample is more likely to be representative of the population, and the sample statistic will be closer to the true population parameter. Statistics provides us with tools to estimate the size of this sampling error, which allows us to say how confident we are in our inferences.


まとめ

今回は、統計的な推測の出発点となる「母集団」と「標本」の関係を見てきました。調べたい対象全体が母集団、そこから選び出して実際に調査する一部が標本であり、全数調査ではなく標本調査を使うのはコスト・時間・実用性の面で利点があるからです。標本から母集団を正しく推測する鍵は、誰もが等しく選ばれる「無作為抽出」によって偏りを避けることであり、それでも生じる標本と母集団のずれが「標本誤差」だということを学びました。

次回は、そもそもデータにはどんな種類があるのかという「データの種類と尺度」を取り上げ、質的データと量的データの見分け方や4つの測定尺度を見ていきます。


日本語訳(全文)

英文を最後まで読み終えてから、答え合わせ用にお使いください。多読の原則として、まずは訳を見ずに英文だけで理解を試みることをおすすめします。

Summary

  • 母集団 対 標本: 母集団(population)とは、調べたり結論を導いたりしたい集団全体のことです(例: ある国のすべての有権者)。標本(sample)は、その母集団から取り出した、扱いやすい小さな部分集合で、実際にデータを集める対象です(例: 世論調査のために選ばれた1,000人の有権者)。
  • 全数調査 対 標本調査: 全数調査(census)は、母集団の一人ひとりすべてからデータを集めます。標本調査(sample survey)は、標本からのみデータを集めます。私たちが標本調査を使うのは、全数調査よりも安く、速く、そして多くの場合より実用的だからです。
  • 無作為抽出の重要性: 母集団について正確な推測を行うには、標本が母集団を代表していなければなりません。母集団のどの一員も等しく選ばれる可能性を持つ無作為抽出(random sampling)は、偏りを避け、代表性のある標本を得るための最良の方法です。
  • 標本誤差: 標本誤差(sampling error)とは、標本から計算した統計量(例: 標本の平均的な意見)と、母集団の真の母数(全員の平均的な意見)との間に自然に生じる差のことです。それは「間違い」ではなく、母集団全体を調べていないことから避けられない結果です。適切に選ばれた大きな標本ほど、標本誤差は小さくなる傾向があります。

全体像: 母集団と標本

統計学では、私たちは大きな個人や物の集団について何かを理解したいと思うことがよくあります。この集団全体のことを母集団(population)と呼びます。母集団は「日本のすべての高校生」のように非常に広いものになることもあれば、「ある工場で生産されたすべての電球」のように、より限定的なものになることもあります。

しかし、母集団の一人ひとり、一つひとつのすべてを調べることは、しばしば困難だったり不可能だったりします。これを全数調査(census)と呼びます。国の国勢調査はこれを行おうとしますが、何年もかかり、莫大な費用がかかる大事業です。たいていの研究やビジネス上の問いにとって、全数調査はまったく現実的ではありません。

では、解決策は何でしょうか。私たちは母集団から、より小さな集団を選んで調べます。この小さな集団のことを標本(sample)と呼びます。そして、この標本からデータを集め、その結果を使って母集団全体についての結論を導きます。このプロセスを標本調査(sample survey)と呼びます。たとえば、テレビ視聴率の会社は、すべての世帯が何を見ているかを監視することはできません。代わりに、数千世帯の標本を選び、その視聴の傾向から国全体の視聴率を推定するのです。

なぜ標本を使うのか

全数調査よりも標本調査を使うことには、いくつかの重要な利点があります。

  • 費用対効果が高い: 数百万人からデータを集めるよりも、数千人から集めるほうがはるかに安上がりです。
  • 時間の節約になる: より小さなデータセットを分析するほうが格段に速いです。これにより、タイムリーな結果を得ることができ、政治の世論調査や経済レポートのようなものにとっては決定的に重要です。
  • 実用性: ときには、データを集める過程そのものが、検査される物を壊してしまうことがあります。たとえば、ある会社が電池の寿命を検査したい場合、すべての電池を使い切るまで試すわけにはいきません。標本を使わなければならないのです。

良い推測の鍵: 無作為抽出

標本を使う目的は、母集団について正確な推測、つまり推量を行うことです。これがうまくいくためには、標本が母集団を代表している必要があります。これは、標本に含まれる人々の特徴が、母集団に含まれる人々の特徴と近く一致しているべきだということを意味します。

では、どうすれば代表性のある標本を得られるのでしょうか。最良の方法は無作為抽出(random sampling)です。単純無作為抽出では、母集団のどの一員も等しく選ばれる可能性を持ちます。これは、全員の名前を巨大な帽子に入れて、そこから名前を引くようなものです。

無作為抽出が決定的に重要なのは、偏り(bias)を避ける助けになるからです。偏りとは、標本を代表性のないものにしてしまう系統的な誤差のことです。たとえば、ある大学の学生の平均身長を知りたいときに、バスケットボール部だけを調査したとすれば、その標本は偏ったものになり、結果も不正確になってしまいます。無作為抽出は、背の高い学生も低い学生も同じように含まれる可能性を保証し、大学全体についてのはるかに正確な姿を与えてくれます。

標本誤差を理解する

完璧な無作為抽出を行ったとしても、標本から得られる結果が、母集団全体から得られる結果とまったく同じになることは、ほとんどありません。ある町の人々の平均年齢を知りたいと想像してください。全住民の真の平均年齢が母数(population parameter)です。さて、あなたは100人の無作為標本を取り、その平均年齢を計算します。これが標本統計量(sample statistic)です。

もし別の100人の無作為標本を取れば、おそらく少し違う平均年齢が得られるでしょう。標本ごとに生じるこの自然なばらつきを標本誤差(sampling error)と呼びます。標本誤差は、あなたの手法における「間違い」ではないということを理解しておくことが大切です。それは、標本が母集団全体のほんの一部にすぎないために生じる、避けられないずれなのです。

標本の大きさを増やすことで、標本誤差を減らすことができます。より大きな標本は母集団を代表している可能性が高く、標本統計量は真の母数により近づきます。統計学は、この標本誤差の大きさを推定する道具を与えてくれ、それによって私たちは自分の推測にどれだけ自信が持てるかを言えるようになるのです。