統計で英語多読 1-3: データの種類と尺度 — 質的データと量的データを見分ける
データの種類と尺度を題材にした英語多読ユニット。約960語の英文と全文日本語訳で、質的・量的データと名義・順序・間隔・比例の4尺度を読みます。
「大人のための英語多読図書館」へようこそ。今回は、カテゴリーで分ける質的データと数値で測る量的データの違い、そして名義・順序・間隔・比例という4つの測定尺度を、英語の文章でたどっていきます。
📊 このユニットの情報 語数: 約960語 / 推定読了時間: 6〜10分 / 難易度: ★★☆☆☆(初中級)
Learning Objectives
みなさんこんにちは。大人のための英語多読図書館へようこそ。
今回は、データ分析の基本中の基本、「データの種類」について学びます。一言で「データ」と言っても、その性質は様々です。統計学では、データを正しく扱うために、まずそのデータがどんな種類のものかを見分けることから始めます。
例えば、「血液型」や「性別」のように、カテゴリーで分けられるデータと、「身長」や「気温」のように、具体的な数値で測れるデータでは、扱い方が全く異なります。前者を質的データ、後者を量的データと呼びます。
さらに、これらのデータはもっと細かく4つの「測定尺度」というものに分類できます。これは、データが持つ情報のレベルを表すものさしのようなものです。この尺度を理解することで、なぜ平均値を計算できるデータとできないデータがあるのか、といった疑問がスッキリ解決します。
この記事では、質的・量的データの違いと、4つの測定尺度(名義、順序、間隔、比例)について、身近な例をたくさん使って解説していきます。この分類は、これから学ぶ様々な分析手法を正しく選択するための、とても大切な土台となります。
それでは今回も多読を楽しんでいきましょう。
Summary
- Qualitative vs. Quantitative Data: Data can be broadly categorized into two types. Qualitative (or Categorical) data represents categories or labels, like gender, nationality, or favorite color. Quantitative (or Numerical) data represents measurable quantities, like height, temperature, or age.
- Four Scales of Measurement: We can further classify data into four levels or scales of measurement, which determine what statistical operations are meaningful.
- Nominal Scale: The simplest level. Data is just a label with no order. Examples: types of cars (Toyota, Honda, Ford), eye color (blue, brown, green). You can count them, but you can't order or average them.
- Ordinal Scale: Data can be ordered or ranked, but the differences between the ranks are not meaningful. Examples: customer satisfaction (Very Satisfied, Satisfied, Dissatisfied), movie ratings (1 to 5 stars). You know one is higher than another, but not by how much.
- Interval Scale: Data is ordered, and the intervals between values are equal and meaningful. However, there is no "true zero." Examples: temperature in Celsius or Fahrenheit, years (e.g., 2023, 2024). You can say 20°C is 10 degrees warmer than 10°C, but you can't say it's twice as hot.
- Ratio Scale: The highest level of measurement. It has all the properties of the interval scale, plus a true zero point. This allows for meaningful ratio comparisons. Examples: height, weight, income, age. You can say someone who is 2 meters tall is twice as tall as someone who is 1 meter tall.
Explanation
First Things First: Identifying Your Data Type
Before we can analyze data, we must first understand what kind of data we have. In statistics, the methods we use depend heavily on the type of data we are working with. The most fundamental way to classify data is to divide it into two main groups: qualitative and quantitative.
-
Qualitative Data: Also known as categorical data, this type of data describes qualities or characteristics. It places individuals or items into categories or groups. Think of it as data you can't do meaningful arithmetic with. For example, your hair color (black, brown, blonde), your nationality (Japanese, American, German), or the type of pet you own (dog, cat, fish) are all qualitative data. We often count the number of items in each category (e.g., 30 people have black hair), but it doesn't make sense to calculate the "average hair color."
-
Quantitative Data: Also known as numerical data, this type represents amounts or counts. It is data that is measured on a numeric scale. Examples include your height in centimeters, the temperature outside in degrees Celsius, or the number of books you've read this year. With this type of data, arithmetic operations like addition, subtraction, and averaging are meaningful.
A Deeper Look: The Four Scales of Measurement
To get even more specific, we can classify data into four levels, or scales, of measurement. These scales tell us about the nature of the information within the values. Moving from the simplest to the most complex, they are: Nominal, Ordinal, Interval, and Ratio.
1. Nominal Scale
This is the most basic level of measurement. The word "nominal" comes from the Latin word for "name." Data on this scale are simply labels or names used to identify different categories. The order of these categories does not matter at all.
- Examples: Blood type (A, B, AB, O), gender (Male, Female, Other), jersey numbers on a sports team.
- What you can do: You can count the frequency of each category (e.g., there are 15 players with blood type A) and find the most common category (the mode). You cannot order them or calculate an average.
2. Ordinal Scale
The word "ordinal" implies order. Data on this scale can be ranked or put in a meaningful order, but the differences between the ranks are not necessarily equal or meaningful. You know the relative position, but not the magnitude of the difference.
- Examples: Survey responses (Strongly Agree, Agree, Neutral, Disagree), educational levels (High School, Bachelor's, Master's, PhD), ranks in a race (1st, 2nd, 3rd).
- What you can do: You can do everything you can with nominal data, plus you can find the median (the middle value) and rank the data. However, you can't say that the difference between "Agree" and "Neutral" is the same as the difference between "Agree" and "Strongly Agree."
3. Interval Scale
Data on the interval scale has a meaningful order, and the intervals (or differences) between the values are equal and meaningful. However, this scale lacks a "true zero" point. A true zero means the complete absence of the thing being measured.
- Examples: Temperature in Celsius or Fahrenheit, calendar years.
- What you can do: You can do everything with ordinal data, plus you can perform addition and subtraction. For example, the difference between 30°C and 20°C is the same as the difference between 20°C and 10°C (both are 10 degrees). However, you cannot make ratio comparisons. You can't say that 20°C is twice as hot as 10°C. This is because 0°C doesn't mean "no heat at all."
4. Ratio Scale
This is the most informative scale of measurement. It has all the properties of the interval scale, but it also includes a true zero point. Because of this absolute zero, we can create meaningful ratios between values.
- Examples: Height, weight, age, income, distance, time.
- What you can do: You can perform all statistical operations: addition, subtraction, multiplication, division, and calculate a wide range of statistics including the mean, median, and mode. You can say that a person who weighs 80 kg is twice as heavy as a person who weighs 40 kg, or that a 4-meter rope is twice as long as a 2-meter rope.
Understanding these scales is crucial because it dictates the kind of analysis you can perform. Using a statistical method designed for ratio data (like calculating the mean) on ordinal data would lead to meaningless and misleading results.
まとめ
今回は、分析を始める前にまず押さえておくべき「データの種類」を見てきました。データは大きく、カテゴリーやラベルで分ける質的データと、数値で測れる量的データに分かれます。さらに細かく見ると、名義・順序・間隔・比例という4つの測定尺度があり、上の尺度に進むほど扱える情報量が増えていきます。名義尺度では平均が計算できないように、尺度の種類によって意味のある分析が決まるという点が、手法を正しく選ぶための土台になります。
次回は、こうしたデータを実際に「見える化」する第一歩として、度数分布表とヒストグラムを使ってデータの分布を可視化する方法を見ていきます。
日本語訳(全文)
英文を最後まで読み終えてから、答え合わせ用にお使いください。多読の原則として、まずは訳を見ずに英文だけで理解を試みることをおすすめします。
Summary
- 質的データ 対 量的データ: データは大きく2種類に分けられます。質的データ(カテゴリカルデータ、質的変数)は、性別・国籍・好きな色のように、カテゴリーやラベルを表します。量的データ(数値データ、量的変数)は、身長・気温・年齢のように、測定可能な量を表します。
- 4つの測定尺度: データはさらに、4つのレベル(尺度)に分類でき、これによってどんな統計的操作が意味を持つかが決まります。
- 名義尺度(Nominal Scale): 最も単純なレベルです。データは順序を持たない、単なるラベルです。例: 車の種類(トヨタ、ホンダ、フォード)、目の色(青、茶、緑)。数えることはできますが、順序づけたり平均したりはできません。
- 順序尺度(Ordinal Scale): データは順序づけたりランクづけたりできますが、ランク間の差には意味がありません。例: 顧客満足度(非常に満足、満足、不満)、映画の評価(1〜5つ星)。どちらが上かは分かりますが、どれだけ上かは分かりません。
- 間隔尺度(Interval Scale): データは順序を持ち、値と値の間隔が等しく意味を持ちます。ただし「真のゼロ」がありません。例: 摂氏・華氏の温度、西暦(例: 2023年、2024年)。20°Cは10°Cより10度暖かいと言えますが、2倍暑いとは言えません。
- 比例尺度(Ratio Scale): 最も高いレベルの測定尺度です。間隔尺度のすべての性質に加えて、真のゼロ点を持ちます。これにより、意味のある比の比較ができます。例: 身長、体重、収入、年齢。身長2メートルの人は1メートルの人の2倍背が高い、と言えます。
まず最初に: 自分のデータの種類を見分ける
データを分析できるようになる前に、まず自分がどんな種類のデータを持っているのかを理解しなければなりません。統計学では、私たちが使う手法は、扱っているデータの種類に大きく左右されます。データを分類する最も基本的な方法は、データを2つの大きなグループ、すなわち質的データと量的データに分けることです。
-
質的データ(質的変数): カテゴリカルデータとも呼ばれ、この種類のデータは性質や特徴を表します。個人や物をカテゴリーやグループに振り分けます。意味のある計算ができないデータだと考えてください。たとえば、髪の色(黒、茶、金)、国籍(日本人、アメリカ人、ドイツ人)、飼っているペットの種類(犬、猫、魚)は、すべて質的データです。私たちは各カテゴリーに含まれる数を数えることがよくありますが(例: 30人が黒い髪をしている)、「平均的な髪の色」を計算することには意味がありません。
-
量的データ(量的変数): 数値データとも呼ばれ、この種類は量や数を表します。数値の尺度で測定されるデータです。例としては、センチメートルで表した身長、摂氏で表した外気温、今年あなたが読んだ本の冊数などがあります。この種類のデータでは、足し算・引き算・平均といった算術操作が意味を持ちます。
より深く見る: 4つの測定尺度
さらに具体的にするために、データを4つのレベル(尺度)の測定に分類できます。これらの尺度は、値が持つ情報の性質について教えてくれます。最も単純なものから最も複雑なものへと並べると、それらは名義尺度、順序尺度、間隔尺度、比例尺度です。
1. 名義尺度(Nominal Scale)
これは最も基本的なレベルの測定です。「nominal(名義の)」という言葉は、ラテン語で「名前」を意味する語に由来します。この尺度のデータは、異なるカテゴリーを識別するために使われる、単なるラベルや名前です。これらのカテゴリーの順序はまったく問題になりません。
- 例: 血液型(A、B、AB、O)、性別(男性、女性、その他)、スポーツチームの背番号。
- できること: 各カテゴリーの頻度を数えること(例: 血液型Aの選手が15人いる)や、最も多いカテゴリー(最頻値)を見つけることができます。順序づけたり平均を計算したりはできません。
2. 順序尺度(Ordinal Scale)
「ordinal(順序の)」という言葉は順序を含意します。この尺度のデータは、ランクづけたり意味のある順序に並べたりできますが、ランク間の差は必ずしも等しかったり意味を持ったりするわけではありません。相対的な位置は分かりますが、差の大きさは分かりません。
- 例: アンケートの回答(強く同意、同意、どちらでもない、不同意)、学歴(高校、学士、修士、博士)、レースの順位(1位、2位、3位)。
- できること: 名義データでできることすべてに加えて、中央値(真ん中の値)を求めたり、データをランクづけたりできます。ただし、「同意」と「どちらでもない」の差が、「同意」と「強く同意」の差と同じだとは言えません。
3. 間隔尺度(Interval Scale)
間隔尺度のデータは意味のある順序を持ち、値と値の間隔(差)が等しく意味を持ちます。ただし、この尺度には「真のゼロ」点がありません。真のゼロとは、測定されているものが完全に存在しないことを意味します。
- 例: 摂氏または華氏の温度、暦の上の年。
- できること: 順序データでできることすべてに加えて、足し算と引き算ができます。たとえば、30°Cと20°Cの差は、20°Cと10°Cの差と同じです(どちらも10度)。しかし、比の比較はできません。20°Cが10°Cの2倍暑いとは言えません。これは、0°Cが「まったく熱がない」ことを意味しないからです。
4. 比例尺度(Ratio Scale)
これは最も情報量の多い測定尺度です。間隔尺度のすべての性質を持ちますが、それに加えて真のゼロ点も含みます。この絶対的なゼロがあるおかげで、値と値の間に意味のある比をつくることができます。
- 例: 身長、体重、年齢、収入、距離、時間。
- できること: 足し算、引き算、かけ算、割り算といったすべての統計的操作を行うことができ、平均・中央値・最頻値を含む幅広い統計量を計算できます。体重80kgの人は40kgの人の2倍重い、とか、4メートルのロープは2メートルのロープの2倍長い、と言うことができます。
これらの尺度を理解することは決定的に重要です。なぜなら、それによってあなたが行える分析の種類が決まるからです。比例データ向けに設計された統計手法(平均の計算など)を順序データに使ってしまうと、無意味で誤解を招く結果につながってしまいます。