Tag Archives: technology

Visual Representation of Complex Data Using R: The tm and wordcloud Packages: Part One Revisited and Updated


I am uploading this “revision of the Revisited and updated” post in a plain text format to better facilliate the use of RSS feed readers to access the article. Links to the PDF version of this posting and the SOTU text file are at the bottom of the my previous post.

After a hiatus during which I retired from a teaching career I am returning to a project that I started nearly five years ago. I have published on this site many tutorials using R analysis to examine large data sets with many variables. My last post on this was part one of a project that involves large data sets composed of text data such as speeches. Virtually any collection of text can be processed using R and numerous R packages have been developed to generate various types of statistical analysis such word frequency counts and correlations among word frequencies. These statistics help enhance the information that is presented visually in the wordcloud graphic generated by the packages discussed in the Part One tutorial.

In this updated version of the Part One tutorial I will review the use of these tools using updated R packages and an analysis using text of the 2026 State of the Union Address. As indicated in Part One, full texts of State of the Union Addresses are available in several formats that can be downloaded. These cover Presidents from George W. Bush to the current term of Donald Trump, and can be found at https://www.govinfo.gov/features/state-of-the-union

There are other sources for SOTU texts but the government website allows quick selection of a file and one click to download the file. I simply save the selected file into a Windows Notepad text file. Make sure when you view the downloaded file that it is in an 80 column format and not formatted to fit a small screen such as a cellphone or tablet. I have found at least a couple of these files that are in the latter format and are not read correctly by the R functions presented in this paper.

The 2026 address was formatted originally in a small screen format. I found a wide-column transcript that I have used in this analysis. A link to the file is found on this website at the end of the article.

This section contains the updated R code that was presented in part one of this project. As indicated above, the resulting wordcloud is from the 2026 State of the Union. Additional code is presented in this section to calculate and print a word frequency table and to print a matrix of word association coefficients for the most used words.

Before diving into the code make sure that you have installed the latest version of the R console or an IDE such as RStudio. As I have stated earlier, I highly recommend the use of RStudio as a development environment and all of the projects presented on this website were coded and refined on this platform. As is the case with any IDE there is a learning curve, but if you are more than a casual programmer the learning process is well worth the effort. If you are using the R console and want to make sure all packages are up to date you can use the following code
#################################################################################script to access and install the newest version of the R console
###############################################################################
#the console should be updated from the Windows R console not from RStudio
#use the following code from the console
################################################################################
install.packages(“installr”) #download installer
library(installr) #move installer to library
updateR() #run the update; follow prompts as needed
################################################################################

The updated wordcloud code as shown below has been extensively commented as to the purpose of the sections within the program structure. For the sake of brevity I will not repeat the explanations contained in the original posting of this document. You can refer to that document on this website if necessary. I should also point out that you can copy all or parts of the code as presented here and paste it into the R console or RStudio code section. It should run properly. I tried to present the code in sections delimited by comments that allow segments to run individually for debugging purposes. As noted in the comments, make sure you change the table function argument section to point to the location of your downloaded text file. The complete code is seen below.


#NOTE:THIS SECTION HAS BEEN UPDATED WITH TRUMP SOTU 2026 DATA
#this is an updated version of Wordcloud1 that will include
#analysis and graphic representation of word counts and association coefficients
#for a single SOTU, Trump 2026
####################################################
#Load required packages
#####################################################
install.packages(“tm”) #processes data
install.packages(“wordcloud”) #creates visual plot
install.packages(“tidyverse”) #graphics utilities
install.packages(“readr”) #to load text files
install.packages(“RColorBrewer”) #for color graphics
#####################################################
#####################################################
#Load and view raw 2026 SOTU text file ‘sotu26.txt’
#Use RStudio Import Dataset tab or code below
#NOTE: BE SURE TO CHANGE THE read_table ARGUMENT TO
#POINT TO THE LOCATION OF THE FILE YOU ARE USING
#####################################################
library(readr)
statu26 <- read_table(“e:/dellfiles/sotu26.txt”, col_names = FALSE) #make sure this points to your file location!
#View(statu26)
####################################################
#Take raw text file statu26 and convert to corpus format named docs26
#####################################################
library(tm)
docs26 <- Corpus(VectorSource(statu26))
####################################################
#Clean punctuation, stopwords, white space
#Three passes create corpus vector source from original file
#A corpus is a collection of text
#One of the functions removes ‘stopwords’, words that are pre-defined
####################################################
library(tm)
library(wordcloud)
data(docs26)
docs26 <- tm_map(docs26, function(x)removeWords(x,stopwords()))
docs26 <- tm_map(docs26,removePunctuation) #remove punctuation
################################################################
#remove stopwords using with ‘en’ or ‘SMART’ criteria
docs26 <- tm_map(docs26,removeWords,stopwords(“SMART”))
###############################################################
docs26 <- tm_map(docs26,stripWhitespace) #remove white space
####################################################
#Cleaned corpus is now formatted into text document matrix
#Then frequency count done for each word in matrix
#dmat <-create matrix; dval <-sort; dframe <-count word frequencies
###################################################
docmat <- TermDocumentMatrix(docs26)
dmat <- as.matrix(docmat)
dval <- sort(rowSums(dmat),decreasing=TRUE)
dframe <- data.frame(word=names(dval),freq=dval)
####################################################
#Final step is to use wordcloud to generate graphics
#There are a number of options that can be set
#Use RColorBrewer to generate a color wordcloud
####################################################
library(RColorBrewer)
set.seed(1234) #use if random.color=TRUE
par(bg=”white”) #background color
wordcloud(dframe$word,dframe$freq,colors=brewer.pal(8,”Set1″),random.order=FALSE,scale=c(2.75,0.35),min.freq=2,max.words=150,rot.per=0.35)
###################################################
####################################################
# Add a title above the plot
mtext(“Donald Trump SOTU 2026”, side = 3, line = 2, cex = 1.5)
##############################################################################
#this section contains code for printing word freq tables;
#word assoc analysis
#########################################################
#code to print 15 most used words with freq
####################################################
#for 2026 address
head(dframe,15)
######################################################
#code to find associations among specified terms
#drops assoc < .25
#initial setup formost frequent words; used between 10 and 25 times
#####################################################
#for 2026 address
findAssocs(docmat,terms = c(“and”,”but”,”budget”,”know”,”economy”,”health”,”plan”,”american”,”energy”,”people”,”care”),corlimit = .25)
######################################################


The output from the code will produce the wordcloud graphic display, a list of the most frequently used words in the document, and a matrix that displays association strengths equal to or greater than .25 for eleven of the most frequently used words in the document. A specific analysis of word frequencies and associations will be the topic of a future paper.

Shown below is the wordcloud graphic, the table of most frequently used words, and a portion of the association matrix generated by the code. I have only shown the portion of the matrix using the four most frequently used words.


Word Frequencies

word freq
and and 25
plan plan 16
but but 15
economy economy 14
now now 13
budget budget 12
health health 12
american american 12
education education 11
energy energy 11
people people 10
care care 10
country country 9
work work 9
america america 9



The association matrix:

$and
along already cmadam first now over second
0.85 0.85 0.85 0.85 0.85 0.85 0.85
speaking still thanks these third the this
0.85 0.85 0.85 0.85 0.85 0.79 0.73
for finally because there you its meet
0.71 0.64 0.50 0.44 0.44 0.41 0.29
bay career claims classroom commerce conscience earned
0.29 0.29 0.29 0.29 0.29 0.29 0.29
fallen ignore import industries ineffective layoffs marketbased
0.29 0.29 0.29 0.29 0.29 0.29 0.29
member message pad perseveres piles pride retooled
0.29 0.29 0.29 0.29 0.29 0.29 0.29
sacrifice succeed warera wealthy
0.29 0.29 0.29 0.29

$but
along already cmadam first over second
0.83 0.83 0.83 0.83 0.83 0.83
speaking still thanks these third now
0.83 0.83 0.83 0.83 0.83 0.81
the this for there finally because
0.81 0.65 0.64 0.64 0.53 0.43
you its agribusiness anger breaks cities
0.43 0.35 0.30 0.30 0.30 0.30
compete csomewhere lower messes pakistan percent
0.30 0.30 0.30 0.30 0.30 0.30
presidents pushed readying reality rebuilt set
0.30 0.30 0.30 0.30 0.30 0.30
transformation view
0.30 0.30

$plan
bank stand held full lose president
0.44 0.42 0.42 0.42 0.42 0.42
weakened builds challenge cmr intend preserve
0.42 0.41 0.41 0.41 0.41 0.41
progress proud respond tells creating cvice
0.41 0.41 0.41 0.41 0.41 0.41
deeds flow friends greensburg honest illusions
0.41 0.41 0.41 0.41 0.41 0.41
ive leonard outstanding prescription speak stopped
0.41 0.41 0.41 0.41 0.41 0.41
strain transform tysheoma boldly candidly doors
0.41 0.41 0.41 0.41 0.41 0.41
join men opinions recognizes showing stability
0.41 0.41 0.41 0.41 0.41 0.41
summit teachers team town save understand
0.41 0.41 0.41 0.41 0.35 0.34
asked recovery
0.30 0.28

$economy
assistance laughter abroad carry comforted costs
0.51 0.51 0.50 0.50 0.50 0.50
cunited cynical direct dramatically gains neighbors
0.50 0.50 0.50 0.50 0.50 0.50
oversight planet receive strategy taking teach
0.50 0.50 0.50 0.50 0.50 0.50
tornado troops auto industry energy act
0.50 0.50 0.37 0.37 0.32 0.30
held hold promise countries government ability
0.30 0.30 0.30 0.30 0.30 0.30
expanded build family social
0.30 0.30 0.30 0.30

In part two of this project I will expand the code to generate comparison clouds and associated statistics for two or more text files. Once again I will use State of the Union texts as the data for analysis. It is my hope that these R tutorials and examples have provided readers with some useful information and working models that can be adapted to their own research projects.

All R programming for this project was done using RStudio 2026.07.0+139 “Pacific Dogwood.”
This PDF document was produced using TeXstudio 4.9.5 (git 4.9.5)
Using Qt Version 6.11.0, compiled with Qt 6.11.0 R.
R, RStudio, and TeXstudio are free, open-source software.
Author: Douglas M. Wiig 7/31/2026
https://dmwiig.net

← Back

Thank you for your response. ✨