where are the x-men?

project source process presentation significance

project

For my midterm project, I used a subset (non-fictional, terrestrial cities) of the dataset of settings in the many issues of the Claremont run and mapped the locations and frequency of each city. Size of each circle correlates to how often it is a setting of an issue.



sources

The data used in this project is drawn from the "Issue Setting" dataset from the digital humanities project "The Claremont Run", which has compiled data from Chris Claremont’s run of the Uncanny X-men comic series. The subset of data contains each issue’s number, setting (whether real or fictional) , time (ie, past, future, or present) and other relevant notes (ie, "in a dream").


processes & data cleaning

For me to be able to efficiently create a useful map of the real-world settings, the data needed significant filtering. First, I used OpenRefine to remove redundant information and create clusters of data. Second, I removed any non-terrestrial or fictional settings. Several issues were set in space, which for obvious reasons I could not map in ArcGIS. Similarly, many issues were set in fictional countries or cities. Some fictional cities set in real countries were specific enough that I was able to map an approximate location, but I filtered out data points of issues set in entirely fictional countries that I wasn’t going to be able to map. Third, I manually changed each data point to list just the location of the city in which in was set, so I would be able to track the popularity on my map of, for example, "Paris, France", instead of "Palais de Justice, Paris, France". Fourth, I created a PivotTable in Excel to add an additional "count" value, so that I could organize my data by how often each city appeared, rather than just the location of each issue in order chronologically. Finally, having narrowed the dataset down to 30 locations and their counts, I used Google Maps to record the coordinates of each location, so that I could export the dataset into ArcGIS. I was originally going to have ArcGIS just use the name of each city, but that resulted in some issues with certain locations, and, due to the relatively small size of the dataset by this point, it was more efficient to manually enter coordinates than to troubleshoot the issue.

After the data was in ArcGIS, I just had the count value correspond to the size of each data point and added some pop-up labels with the name of each location and its frequency. By mapping the "count", I was able to make a map that demonstrates which settings were frequently visited during this run of the Uncanny X-Men comics, and which were not.


presentation

When creating this website to display the information, I wanted the design to relate to the subject of the data itself. Since the subject of the data is pretty fun and lighthearted, I used brighter colors and a fun background, and made most elements red, white, and blue because of their association with American comics. I chose to make a pretty straightforward one-page layout, but since one page can be a lot to scroll through (and I like making buttons) I added links at the top of the page to take the user to each section of the page using "id", like a table of contents.


significance

By applying tools like ArcGIS not just to real-world locations and data but instead to data about settings of fiction, we can make some insights about how the media in question often operates. One of the results from this representation of this data, for example, is that the issues in this run are most frequently set in American cities, particularly New York, and other frequently storyline-relevant locations, like Australia. Since the comics were produced and sold in America, it’s easy to understand why this would be true. In general, mapping fictional settings like in this project allows us to draw conclusions or arguments about fiction and popular media that: a.) use a much larger volume of data than would be useful in a paper or written argument, and b.) are more accessible and easily readable to a larger audience. The latter is particularly important when analyzing media in popular culture like comics-- if the subject matter is for a general audience, why should its analysis not be?

Projects like these are only made possible by the field of Digital Humanities– on the one hand, traditional data science projects might not venture into this kind of (more humanities-focused) subject matter; on the other hand, more traditional humanities scholars would not be able to use the same volume of data as used in this project in something like a paper, without the use of digital humanities & data analysis tools.


back to top