A Text Mining Odyssey: Using SV’s new AI Text Categoriser to analyse reactions to Christopher Nolan’s summer blockbuster

For decades, data analysts have struggled with problem of transforming unstructured data into structured information. Nowhere is this more evident than in how researchers have dealt with fields comprised of open text. Historically, only the most intrepid analysts have attempted to understand or categorise the content of these variables. To do so, they have invariably resorted to painstaking and time-consuming approaches such as manual coding schemes or complex rule-driven classification methods.

The reason they have gone to such efforts is simply because the potential insights that open text variables offer are powerful and valuable. Despite this, for the majority of analysts, the effort required to unlock any real value from text data is so burdensome, that they are simply ignored. However, the advent of AI and large language models has changed that. So much so, that for most of these models, it is a relatively trivial task to ask them to do what previous technologies found extremely difficult: read a text response, understand the content and categorise the subject matter.

The Watsonx.ai platform provides users with access to three third-party language models licensed by IBM: Meta Llama 3.3 70B Instruct, Meta Llama 4 Maverick, and Mistral Small 3.1 24B Instruct. Moreover, IBM states that prompts and outputs submitted through its Watsonx.ai platform are private and that it does not use LLM model inputs or outputs to improve the models and does not monitor or log the inputs or outputs.

To that end, Smart Vision Europe has developed a simple plug and play extension that enables users of IBM SPSS Statistics to access these capabilities. Using the SV Watsonx.ai Text Categoriser, users can automatically uncover topics and themes in text data, categorise each response and reveal deep insights.

To illustrate this, we collected comments from 100 readers of The Guardian about Christopher Nolan’s summer blockbuster, The Odyssey and used the Text Categoriser to summarise the main themes in their reactions to the film. Using the program’s Discovery mode, we requested that it analyse the text and return a maximum of 9 categories that illustrated the main topics discussed. In fact, we did this using all three of the language models and compared the results which are shown below.

As can be seen from the comparison table, Discovery mode instantly revealed the core themes of the discussion. Despite the fact that each language model chose a different number of underlying topics to summarise the data (ranging from 5 topics to 7), there are considerable overlaps between the results. All three LLM models identified themes related to the screenplay adaption, sound quality, visuals and cinematography as well as casting/character portrayal. This procedure alone delivered considerable insight as to the which aspects of the film provoked most discussion.

But we can go further, as part of the purpose of the Discovery mode is to provide the user with a basis on which to categorise the responses themselves and in doing so, drive deeper analysis. It does this by creating a Topic Map file. This simple file can then be used to ensure that the categorisation process only uses a list of specified topics as the basis for creating categories. In fact, users can define their own topic file, and after experimenting with Discovery mode and merging on the results from different runs, we decided to create a custom topic file comprised of the following topic categories:

  1. Adaptation and Faithfulness
  2. Character Portrayal
  3. Sound and Music
  4. Visuals and Cinematography
  5. Pacing and Length
  6. Emotional Resonance
  7. Casting and Representation
  8. Nolan’s Direction

Having defined our topic file, we then switched to the Categorisation mode. Executing this this mode allowed us to categorise the entire dataset and see the how frequently each topic category was mentioned.

The chart immediately showed that the most common topic discussed was character portrayal. Typically, this referred to the portrayal of Odysseus and how the film used supporting characters. In total, 43% of the readers in the sample mentioned this aspect of the film.

  • “John Leguizamo as Eumaeus very moving.”
  • “One thing that maybe could have been done a bit differently was the wimpyness of Telemachus”
  • “Calypso was the one character that didn’t really resonate with me”

In contrast, the least common topic was Casting and Representation (17%). Here the commenters referred to issues of diversity and the suitability of actors cast in in different roles.

  • “Zendaya winning supporting actress would be a joke.”
  • “I saw a lot of ridicule and uproar at the casting choices”
  • “Provocative casting, like Clytemnestra and Helena being twin sisters”

In fact, we can go further, as this mode not only categorised the data based on the 8 topics in the topic map file it also categorised them based on sentiment. This means the procedure created an additional 24 variables as each topic was categorised by whether the reader referred to it with a positive, negative or neutral sentiment.  Doing this added a key dimension to the analysis.

The chart below shows the topic breakdown by sentiment. What this reveals is that by far the most positively evaluated topic was Visuals and Cinematography.  In fact, 35% of the respondents praised this aspect of the film. In contrast, 23.7% of the readers specifically criticised the sound and/or the music with many references to loudness of and difficulty with hearing the actors’ dialogue.

  • “Nolan’s sound seems always to manage to be very loud while making it difficult to hear what people are saying.”
  • “The noise levels far too much at times.”
  • “Totally agree regarding the eardrum shredding soundtrack.”

Neutral topic classifications often appeared where readers tried to balance good and bad aspects of a topic or offered lukewarm praise. We can see this in the neutral classifications for Nolan’s Direction where comments often weighed up the choices the director made and compared them to other films.

  • “Is it just me? Nolan films tend to leave me impressed rather than entranced”.
  • “For me, the script is not terrible and the acting not terrible, but it’s as if Nolan doesn’t really set much store in either.”
  • “Whilst I enjoyed Nolan’s movie, whenever the setting was in Ithaca I found myself comparing it to The Return”

The Text Categoriser also examined how the topics were related to each other. Among those who offered criticism of the film, there appeared to be some evidence of correlation between negative reviews of key aspects. For example, those who expressed negativity about casting and representation also tended to be negative about adaptation and faithfulness (Pearson correlation 0.341). Tellingly, people those who were critical of the film’s emotional resonance were much more likely to be critical of Nolan’s direction (Pearson correlation 0.339) than they were of character portrayal (Pearson correlation 0.077) or casting and representation (Pearson correlation 0.051).

One aspect of the Text Categoriser’s analytical abilities that we weren’t able to test out, was how these uncovered topics related to a key target variable. Had the respondents simply been asked whether or not they recommended the film to others, their binary Yes/No responses could have formed the basis of deeper analysis. This is because the Text Categoriser is able to automatically perform a key driver analysis where it compares the impact of each categorised topic upon an outcome such as likelihood to recommend. In doing so, the resulting tables and charts help analysts to measure how much the various positive and negative topics drive the recommendation likelihood up or down. 

If you have SPSS Statistics, you can try out the SV Watsonx.ai Text Categoriser yourself for free. Simply download the Demo version of the tool (including a sample dataset) from here and let us know how you get on.

If you’d like to learn more take a look at our detailed guide to setting up and using the Text Categoriser.