英語多読

統計で英語多読 5-1: 標本抽出法 — 良い推測のためのサンプリング

標本抽出法を題材にした英語多読ユニット。約790語の英文と全文日本語訳で、単純無作為・層化・クラスター・多段抽出の使い分けをたどります。

大人のための英語多読図書館」へようこそ。今回は、全体の一部である「標本」をどう選べば信頼できる推測ができるのか、単純無作為抽出から多段抽出までの代表的なサンプリング手法を、英語の文章でたどっていきます。

📊 このユニットの情報 語数: 約790語 / 推定読了時間: 5〜8分 / 難易度: ★★☆☆☆(初中級)

Learning Objectives

みなさんこんにちは。大人のための英語多読図書館へようこそ。

国全体の平均年収や、ある製品に対する満足度を知りたいとき、対象者全員にアンケートを取るのは現実的ではありませんよね。時間もコストもかかりすぎてしまいます。そこで統計学では、全体の一部である「標本(サンプル)」をうまく選び出し、そこから全体(母集団)の姿を推測するというアプローチを取ります。

しかし、この「サンプルの選び方」が非常に重要です。偏ったサンプルを選んでしまうと、推測結果も偏ったものになってしまいます。今回は、信頼できる推測を行うための科学的なサンプルの選び方、「標本抽出法」について学びます。単純無作為抽出、層化抽出、クラスター抽出など、様々な方法のメリット・デメリットを理解していきましょう。

それでは今回も多読を楽しんでいきましょう。

Summary

  • Sampling is the process of selecting a subset of individuals from a larger population to make inferences about the whole population.
  • A census, which surveys the entire population, is often impractical due to high costs and time consumption.
  • Simple Random Sampling (SRS) gives every member of the population an equal chance of being selected. It is unbiased but can be difficult to implement for large populations.
  • Stratified Sampling involves dividing the population into subgroups (strata) based on shared characteristics and then performing simple random sampling within each subgroup. This ensures representation from all key groups.
  • Cluster Sampling divides the population into clusters (often based on geography), randomly selects some of these clusters, and then samples all individuals within the selected clusters. It is cost-effective but may be less precise.
  • Multi-stage Sampling is a more complex method that combines different sampling techniques, often used in large-scale national surveys.
  • Choosing the right sampling method is crucial for obtaining representative data and making reliable statistical inferences.

Explanation

Why Do We Need Sampling?

Imagine you want to know the average height of all adults in Japan. It would be impossible to measure every single person. This is where sampling comes in. Instead of studying the entire population (all adults in Japan), we select a smaller group, called a sample. By studying the sample carefully, we can make an educated guess, or an inference, about the entire population. The goal is to choose a sample that is a mini-version of the population. If our sample is representative, our conclusions will be accurate. If it's biased, our conclusions will be wrong. Let's explore the main methods for selecting a good sample.

Simple Random Sampling (SRS)

This is the most straightforward sampling method. In simple random sampling, every individual in the population has an equal chance of being chosen. It's like putting everyone's name into a giant hat and drawing names at random.

  • Advantage: The biggest advantage of SRS is that it's highly unbiased. If done correctly, the sample is very likely to be representative of the population.
  • Disadvantage: It can be hard to execute in practice. You need a complete list of every single person in the population (called a sampling frame), which is often unavailable. Also, by pure chance, you might end up with a sample that doesn't represent smaller subgroups well.

Stratified Sampling

Suppose you are conducting a survey about political opinions. You know that age is a very important factor. If you use SRS, you might, by chance, get a sample with mostly young people. To avoid this, you can use stratified sampling.

In this method, you first divide the population into different subgroups, or strata, based on a shared characteristic like age, gender, or income level. Then, you perform simple random sampling within each stratum. For example, you could divide the population into age groups (20-29, 30-39, 40-49, etc.) and then randomly select a proportional number of people from each group.

  • Advantage: Stratified sampling guarantees that your sample will be representative with respect to the characteristics you used to create the strata. This often leads to more precise estimates than SRS.
  • Disadvantage: You need to have prior knowledge about the population to be able to divide it into strata. It can also be more complex to organize.

Cluster Sampling

Now, imagine you want to survey students from all high schools in a large city. Visiting randomly selected students across hundreds of schools would be very expensive and time-consuming. Instead, you could use cluster sampling.

In cluster sampling, you divide the population into groups, or clusters, which are often geographical (e.g., cities, school districts). You then randomly select a certain number of these clusters. Finally, you survey every single individual within the selected clusters.

  • Advantage: This method is much more practical and cost-effective, especially when the population is spread over a wide geographic area.
  • Disadvantage: The results can be less precise than SRS or stratified sampling if the clusters themselves are not very representative of the overall population. For example, if you randomly select schools from only wealthy neighborhoods, your results will be biased.

Multi-stage Sampling

Multi-stage sampling takes this process a step further. It involves several stages of sampling. For instance, in a national survey, you might: 1. Stage 1: Randomly select a number of prefectures (clusters). 2. Stage 2: Within each selected prefecture, randomly select a number of cities or towns (clusters). 3. Stage 3: Within each selected city, randomly select a number of neighborhoods (clusters). 4. Stage 4: Within each selected neighborhood, randomly select a number of households to survey.

This complex method is widely used in large-scale, real-world research because it balances precision with practicality. The choice of sampling method always depends on the research question, budget, and available resources.


まとめ

今回は、母集団全体を調べる代わりに一部である標本を選び、そこから全体を推測する「標本抽出法」を見てきました。全員に当選確率が等しい単純無作為抽出、属性ごとに層を分けてから抽出する層化抽出、地理的なまとまりを選ぶクラスター抽出、そしてそれらを段階的に組み合わせる多段抽出。それぞれに精度とコストの面で得意・不得意があり、目的や予算に応じて使い分けることが、信頼できる推測の出発点になります。

次回は、選び出した標本データから母集団の値を「一点」で言い当てる「点推定」を取り上げ、良い推定量が満たすべき性質と、標本分散をn-1で割る理由を見ていきます。


日本語訳(全文)

英文を最後まで読み終えてから、答え合わせ用にお使いください。多読の原則として、まずは訳を見ずに英文だけで理解を試みることをおすすめします。

Summary

  • 標本抽出(サンプリング)とは、母集団全体について推測を行うために、より大きな母集団の中から一部の個体を選び出す手続きのことです。
  • 母集団全体を調査する全数調査(センサス)は、コストと時間がかかりすぎるため、しばしば現実的ではありません。
  • 単純無作為抽出(SRS)は、母集団のすべての構成員に選ばれる確率を等しく与えます。偏りがありませんが、大きな母集団では実施が難しいことがあります。
  • 層化抽出は、共通の特性にもとづいて母集団を部分集団(層)に分け、それぞれの層の中で単純無作為抽出を行う方法です。これにより、すべての重要なグループから代表を確保できます。
  • クラスター抽出は、母集団を(多くは地理にもとづく)クラスターに分け、そのうちのいくつかを無作為に選び、選ばれたクラスター内のすべての個体を調査する方法です。費用対効果は高いものの、精度は劣ることがあります。
  • 多段抽出は、異なる抽出技法を組み合わせた、より複雑な方法で、大規模な全国調査でよく用いられます。
  • 適切な抽出法を選ぶことは、代表性のあるデータを得て、信頼できる統計的推測を行うためにきわめて重要です。

なぜ標本抽出が必要なのか

日本の成人全員の平均身長を知りたいと想像してみてください。一人ひとりを測定するのは不可能でしょう。ここで標本抽出の出番です。母集団全体(日本の成人全員)を調べる代わりに、標本と呼ばれる、より小さな集団を選び出します。標本を注意深く調べることで、母集団全体について根拠のある推測、すなわち推論を行うことができます。目標は、母集団のミニチュア版になっている標本を選ぶことです。標本が代表性を持っていれば、結論は正確になります。偏っていれば、結論は誤ったものになります。それでは、良い標本を選ぶための主な方法を見ていきましょう。

単純無作為抽出(SRS)

これは最も素直な抽出法です。単純無作為抽出では、母集団のすべての個体が選ばれる確率を等しく持ちます。全員の名前を巨大な帽子に入れて、無作為に名前を引くようなものです。

  • 長所: SRSの最大の長所は、偏りが非常に小さいことです。正しく行えば、標本は母集団を代表している可能性が高くなります。
  • 短所: 実際には実施が難しいことがあります。母集団に含まれるすべての人の完全なリスト(抽出枠と呼ばれます)が必要ですが、これはしばしば入手できません。また、まったくの偶然によって、小さな部分集団をうまく代表できていない標本になってしまうこともあります。

層化抽出

政治的な意見に関する調査を行っているとしましょう。年齢が非常に重要な要因だと分かっているとします。SRSを使うと、偶然によって若い人ばかりの標本になってしまうかもしれません。これを避けるために、層化抽出を使うことができます。

この方法では、まず年齢・性別・所得水準といった共通の特性にもとづいて、母集団を異なる部分集団、すなわちに分けます。次に、それぞれの層の中で単純無作為抽出を行います。たとえば、母集団を年齢層(20〜29歳、30〜39歳、40〜49歳など)に分け、各グループから比率に応じた人数を無作為に選ぶ、といった具合です。

  • 長所: 層化抽出は、層を作るために使った特性に関して、標本が代表性を持つことを保証します。これにより、SRSよりも精度の高い推定につながることがよくあります。
  • 短所: 母集団を層に分けられるよう、あらかじめ母集団についての知識を持っている必要があります。また、組み立てがより複雑になることもあります。

クラスター抽出

今度は、大都市にあるすべての高校の生徒を調査したいと想像してみてください。何百もの学校にまたがって無作為に選ばれた生徒を訪ねるのは、非常に費用と時間がかかるでしょう。その代わりに、クラスター抽出を使うことができます。

クラスター抽出では、母集団を、しばしば地理的な(たとえば市や学区などの)グループ、すなわちクラスターに分けます。次に、そのクラスターのうち一定数を無作為に選びます。最後に、選ばれたクラスター内のすべての個体を調査します。

  • 長所: この方法は、特に母集団が広い地理的範囲に散らばっている場合、はるかに実用的で費用対効果が高くなります。
  • 短所: クラスターそのものが母集団全体をあまり代表していない場合、結果はSRSや層化抽出よりも精度が劣ることがあります。たとえば、富裕層の地域からだけ学校を無作為に選んでしまえば、結果は偏ってしまいます。

多段抽出

多段抽出は、このプロセスをさらに一歩進めます。これは複数の段階の抽出を含みます。たとえば、全国調査では次のように行うかもしれません。 1. 第1段階: いくつかの都道府県(クラスター)を無作為に選ぶ。 2. 第2段階: 選ばれた各都道府県の中で、いくつかの市町村(クラスター)を無作為に選ぶ。 3. 第3段階: 選ばれた各市の中で、いくつかの地区(クラスター)を無作為に選ぶ。 4. 第4段階: 選ばれた各地区の中で、調査する世帯をいくつか無作為に選ぶ。

この複雑な方法は、精度と実用性のバランスを取れるため、大規模で現実的な研究において広く用いられています。抽出法の選択は、つねに研究上の問い、予算、利用できる資源に左右されます。