Reusable Scatter Plot: Hours vs Exam (color =
Attendance%, radius = Assignments)
Explanation & Analysis (for submission)
Dataset (not Iris): 50 synthetic rows of student performance with columns:
- hours (X), exam (Y), attendance_pct (color), assignments (radius), student (label).
Reusable component:
- The chart is built with a factory function ScatterPlot() exposing setters:
.x(), .y(), .r(), .color(), .width(), .height(), .xLabel(), .yLabel(), .title(), .rScale()
- This makes the plot reusable with any dataset by just changing accessors, not the internals.
How attributes are distributed:
- X vs Y shows an upward trend: more hours -> higher exam scores.
- Color encodes attendance% using thresholds: <=60 (blue), 60–80 (purple), >80 (red).
At similar hours, red points often lie above blue points (attendance seems to boost outcomes).
- Radius encodes assignments; larger circles cluster in higher exam regions, reinforcing that steady work correlates with exam performance.
Why build a reusable scatter plot:
- Consistency: same axes, labeling, and behaviors across datasets and assignments.
- Maintainability: improvements (e.g., tick formatting, titles) apply everywhere the component is reused.
- Flexibility: switch to a new dataset by swapping accessors only; supports dashboards and iterative assignments.
- Efficiency: avoids rewriting boilerplate (scales, axes, tooltips) each time.