Relationships: Cross Tables, Scatter Plots, and Beyond
Data visualization is a cornerstone of data analysis. It allows us to transform raw numbers into clear and informative pictures, helping us understand patterns, trends, and relationships between variables. Let’s get into two fundamental data visualization techniques for exploring relationships between two variables: cross tables and scatter plots. We’ll also shed light on how to identify the data type to choose the most appropriate visualization method.
Categorical Encounters: The Power of Cross Tables
When dealing with categorical variables, also known as nominal variables, cross tables (or contingency tables) reign supreme. You’re a financial scholar managing investment portfolios for three clients with varying risk tolerances. Each client’s portfolio reflects their risk preference, with allocations diversified across stocks, bonds, and real estate.
A cross table would elegantly capture this data, as shown below:
| Investment Type | Investor A | Investor B | Investor C | Total |
| Stocks | $10,000 | $20,000 | $5,000 | $35,000 |
| Bonds | $5,000 | $15,000 | $10,000 | $30,000 |
| Real Estate | $5,000 | $5,000 | $15,000 | $25,000 |
| Total | $20,000 | $40,000 | $30,000 | $90,000 |
This table offers a clear view of how each investor’s portfolio is distributed across asset classes (rows) and the total holdings of each investor (columns). Calculating row and column totals further enhances the analysis. Row totals reveal the total investment in each asset class, while column totals showcase each investor’s overall portfolio breakdown.
Taking visualization a step further, side-by-side bar charts excel at portraying data from cross tables. Each investor’s asset allocation is depicted using bars of distinct colors, allowing for easy comparison of investment strategies across clients and within each client’s portfolio.
Numerical Narratives: The Magic of Scatter Plots
Scatter plots come into play when we have two numerical variables, also known as quantitative variables. Let’s explore the SAT reading and writing scores of 100 students.
Here’s a breakdown of a scatter plot representing this data:
- Data Range: Both reading and writing scores range from 200 to 800.
- Axes: The vertical axis (y-axis) represents writing scores, while the horizontal axis (x-axis) showcases reading scores.
- Data Points: Each of the 100 students is represented by a single dot on the graph, with the position of the dot corresponding to their reading and writing scores.
By analyzing the scatter plot, we can uncover fascinating stories within the data:
- Upward Trend: A general upward trend signifies that students with higher reading scores tend to achieve higher writing scores, and vice versa. This positive correlation makes sense, as both reading and writing skills are often interconnected.
- Central Clustering: A concentration of students in the middle of the graph, with scores between 450 and 550 on both axes, reflects the average score range (around 500) for both reading and writing.
- Exceptional Performers: A cluster of students with exceptionally high scores in both reading and writing highlights the presence of high-achieving individuals who excel in both areas.
- Outliers: Data points that deviate significantly from the overall trend, like Jane’s score (high reading, low writing), are termed outliers. These warrant further investigation as they might indicate unique circumstances or data recording errors.
Understanding how to choose the right visualization technique for categorical and numerical data empowers you to extract valuable insights from your data. Cross tables and scatter plots are fundamental tools in this journey, providing a clear and concise way to explore relationships between variables. Leveraging these techniques, you can transform raw data into a compelling narrative, uncovering hidden patterns and trends that would otherwise remain unseen.
For further information: Crosstab, Contingency tables, Scatter plots


Leave a comment