Skip to main content
100%

Data Analysis — starwars_edges.csv

✓ Published20🌍 Public
TTest
Last edited Oct 3, 2026
Created on Oct 3, 2026

Loads the project dataset starwars_edges.csv (Star Wars character co-occurrence edges), showing a dataset summary, a first-10-rows preview table, and README documentation of every column with type classifications.

Star Wars Interaction Edges Dataset

Dataset Overview

  • Dataset file: starwars_edges.csv
  • Total row count: 1013
  • Total column count: 4

The dataset is a long-form directed edge list describing character interactions in Star Wars films. Each row represents an interaction between a source character and a target character in a particular episode. The same character pair can appear across multiple episodes, so rows do not necessarily represent unique pairs.

The visualization loads the dataset at runtime from /data/starwars_edges.csv. The displayed row count is computed from the loaded data rather than hardcoded.

Column Descriptions

episode

  • Description: The Star Wars episode number in which the interaction occurs.
  • Data type: Ordinal / nominal grouping. Episode values range from 1 through 6. Although the values have an inherent numeric sequence, they are primarily used as episode categories or groups.
  • Missing values and assumptions: The column is parsed as a number. Values are expected to be integers from 1 to 6. No obvious missing values were observed in the sampled rows.

source

  • Description: The name of the character serving as the source of the directed interaction edge.
  • Data type: Nominal. Character names are categorical labels with no inherent ordering.
  • Missing values and assumptions: Values are parsed as strings. Character names may include spaces, punctuation, or dashes. No obvious missing values were observed in the sampled rows.

target

  • Description: The name of the character serving as the target of the directed interaction edge.
  • Data type: Nominal. Character names are categorical labels with no inherent ordering.
  • Missing values and assumptions: Values are parsed as strings. Character names may include spaces, punctuation, or dashes. No obvious missing values were observed in the sampled rows.

value

  • Description: The count of interactions or scenes shared between the source and target character pair in the specified episode.
  • Data type: Quantitative. Arithmetic operations such as sums and comparisons are meaningful for these integer counts.
  • Missing values and assumptions: The column is parsed as an integer count. Values are expected to be non-negative whole numbers. No obvious missing values were observed in the sampled rows.

Data Structure Notes

  • The dataset contains 1013 rows and 4 columns.
  • It is a long-form edge list rather than a table with one row per unique character pair.
  • The same pair may occur in multiple episodes.
  • source and target define a directed relationship, so their order is meaningful.
  • The visualization uses the actual loaded dataset for the row count and preview.