Data Analysis — starwars_edges.csv
Loads and documents the starwars_edges.csv project dataset (1013 rows, 4 columns).
Star Wars Edges — Dataset Explorer
A React + D3 single-page visualization that loads the starwars_edges.csv
project dataset at runtime and renders:
- A dataset summary (name, file size, row count, column count).
- A preview table of the first 10 rows with all column headers.
- This README file, documenting the dataset schema.
Loading
The dataset is loaded at runtime from its project-dataset virtual path:
/data/starwars_edges.csv
Schema
| Column | Type | Description |
|---|---|---|
episode |
Ordinal | Episode number (1–6) |
source |
Nominal | Origin character in interaction |
target |
Nominal | Destination character in interaction |
value |
Quantitative | Weighted interaction strength (integer ≥ 0) |
The dataset is a directed, weighted edge list representing character interactions across all Star Wars films (Episodes I–VI). Each row denotes that character source interacts with character target in a given episode, with value representing the intensity (e.g., number of scenes, lines spoken together, or other normalized proximity metric).
All edge values are non-negative integers. The dataset is sparse: most character pairs appear only once or twice across the full corpus.
Visualization Ideas
Episode-Binned Force-Directed Network
- Question: How do character communities evolve across episodes?
- Fields: episode, source, target, value
- Encoding: Nodes = characters (radius ∝ degree centrality), edges = weighted (opacity/width ∝ value), color = episode (1–6), layout via force-directed.
- Interaction: Toggle episodes on/off; hover to show edge weight and characters; click to highlight ego-network of a selected node.
- Note: Render 6 separate force layouts (one per episode) or animate transitions between episodes.
Aggregated Adjacency Matrix (Heatmap)
- Question: Which character pairs interact most often, and in which episodes?
- Fields: episode, source, target, value
- Encoding: Rows/columns = sorted characters (by total degree), cell = sum of values across episodes, color = intensity (log-scaled), shape = binary presence.
- Interaction: Hover to show (source, target, total weight, episode breakdown); click a cell to filter table to those edges.
- Note: Use hierarchical clustering (e.g., hierarchical edge bundling on matrix) to reorder rows/columns.
Episode-Parallel Arc Diagram
- Question: How do interactions cluster spatially within each episode?
- Fields: episode, source, target, value
- Encoding: Horizontal layout per episode; nodes on x-axis (alphabetical/centralized), arcs over axis (thickness ∝ value), color = episode.
- Interaction: Hover arcs to show edge weight; click episode label to highlight only that arc layer.
- Note: Ideal for visualizing temporal proximity (e.g., sorting by narrative order).
Sankey Flow Diagram
- Question: How do characters flow between episodes in terms of interaction volume?
- Fields: episode, source, target, value
- Encoding: Nodes = characters per episode (left-right across episodes), flows = weighted edges, color = source character group.
- Interaction: Hover flows for value; toggle episodes to isolate segment.
- Note: Require node expansion to avoid excessive edge bundling; consider splitting nodes across episodes.
Chord Diagram (Episode Aggregated)
- Question: What is the pattern of cross-character interaction strength within and across episodes?
- Fields: episode, source, target, value
- Encoding: Circular layout; nodes = top N characters; ribbons = aggregated value between pairs, fill = symmetric weight.
- Interaction: Hover to show source, target, aggregated weight, episode distribution.
- Note: Aggregate across episodes first; use symmetric matrix (sum of A→B + B→A) if direction is not essential.
Small Multiples Timeline (by Episode)
- Question: How does the character interaction network evolve per episode?
- Fields: episode, source, target, value
- Encoding: 6 small force-directed or edge-bundled networks, one per episode; node size ∝ episode-degree; edge width ∝ episode-weight.
- Interaction: Synced brushing across plots; click episode to highlight in table.
- Note: Align node positions (e.g., spring embedder with fixed seed) for comparability.
Character Rank-Order Bar Chart
- Question: Who are the most central characters in each episode?
- Fields: episode, source, target, value
- Encoding: Bar chart with y-axis = character, x-axis = sum(incoming value), bars grouped/stacked by episode.
- Interaction: Sort by total, episode, or alphabetically; hover to show breakdown by episode.
- Note: Use horizontal bars for long character names; invert axis if > 30 characters.
Weight Distribution Histogram
- Question: What is the distribution of interaction strengths across the dataset?
- Fields: value
- Encoding: Histogram bins for value counts; overlay density curve if needed.
- Interaction: Click bin to filter table by edge weight range.
- Note: Use log-scale x-axis due to heavy-tailed distribution.
Episode Heatmap Matrix (Character × Episode)
- Question: Which characters appear in which episodes, and how strongly?
- Fields: episode, source, target, value
- Encoding: Rows = top N characters, columns = episodes, cell = max(incoming+outgoing value) for that character/episode, color intensity.
- Interaction: Hover to show character’s total episode contribution; click to highlight all related edges.
- Note: Normalize per character/episode if scale differences dominate.
Cumulative Weight Timeline
- Question: How do cumulative interaction volumes build across episodes?
- Fields: episode, source, target, value
- Encoding: Line/area chart: x = episode, y = cumulative sum of edge weights, lines = top characters.
- Interaction: Toggle characters; hover for exact cumulative value.
- Note: Could show per-character or aggregated over all characters.
Edge Duration Stacked Bar (Episode-Specific)
- Question: Which character pairs co-appear longest in each episode?
- Fields: episode, source, target, value
- Encoding: Horizontal stacked bars: one row per (source, target) pair, color = episode, segment length = value.
- Interaction: Click pair to highlight all episodes for that edge.
- Note: Limit rows to top 30 pairs to avoid clutter.
Character Trajectory Graph
- Question: How does a character’s influence (degree, weighted in/out) change across episodes?
- Fields: episode, source, target, value
- Encoding: Line chart: x = episode, y = degree/weight metric, separate line per character.
- Interaction: Hover for values; select character to highlight in network views.
- Note: Show both in-degree and out-degree as separate series.
Interactive Edge List Table with Filters
- Question: How can users explore raw edges by custom criteria?
- Fields: episode, source, target, value
- Encoding: Tabular view with column filters (episode dropdown, character search, value slider).
- Interaction: Filter in real time; export selected rows; link to visualizations (e.g., click row to show edge in network).
- Note: Use Debounce on filters; support regex on character names.
Episode Co-occurrence Network (Character × Episode)
- Question: Which episodes share the same core character groups?
- Fields: episode, source, target, value
- Encoding: Bipartite network: one set = characters, other = episodes, edges = character appears (value > 0), weight by sum of incident edge values.
- Interaction: Click episode to show character set; click character to show episode list.
- Note: Project to character-only (if same episode co-occurrence desired) via matrix multiplication.
Radial Edge Bundling (Episode → Characters)
- Question: Do interaction patterns cluster around central figures per episode?
- Fields: episode, source, target, value
- Encoding: Radial layout: center = episode hub, radial spokes = characters (sorted by episode centrality), edges = weighted arcs to other characters.
- Interaction: Expand character by clicking; show tooltip with centrality metrics.
- Note: Simplify by using only top 20% degree characters per episode.
Topological Sorting for Narrative Chronology
- Question: Can we infer character appearance sequence or centrality order across episodes?
- Fields: episode, source, target, value
- Encoding: Time-series edge diagram (x = episode, y = character name), edges = source→target over episodes.
- Interaction: Hover edges to show weight; group characters by episode debut or cumulative weight.
- Note: Order characters by debut episode or total episode participation.
Betweenness Centrality Animated Map
- Question: Which characters act as bridges across episodes?
- Fields: episode, source, target, value
- Encoding: Bar chart of betweenness centrality (computed per episode), bars animated over time; color = betweenness percentile.
- Interaction: Toggle episode; hover to show character name and value.
- Note: Precompute centrality per episode to avoid runtime overhead.
Episode Pair Sankey Flow
- Question: How do character sets overlap between consecutive episodes?
- Encoding: Nodes = characters in episode i, edges to characters in episode i+1, flow = shared edge weight.
- Fields: episode, source, target, value
- Interaction: Hover to show overlap strength; click pair to isolate episode jump.
- Note: Focus on sequential episode pairs (e.g., 1→2, 2→3, etc.).
Weighted Edge Cumulative Curve
- Question: What proportion of total interaction weight comes from top edges?
- Fields: value
- Encoding: Pareto-style curve: x = edges sorted by weight (desc), y = cumulative % of total weight.
- Interaction: Hover to show weight threshold and % coverage; click point to filter table.
- Note: Add reference lines (e.g., 80% of weight from top 10% edges).
Character Interaction Timeline Heatmap
- Question: When do character pairs first/last appear together?
- Fields: episode, source, target, value
- Encoding: Rows = top edge pairs, columns = episodes 1–6, cell = max weight in episode or binary first/last.
- Interaction: Click pair to highlight in network or show all edges for that pair.
- Note: Use binary for first appearance; color by value for repeated interactions.