Outbound Wiki

PDF

Spam and Educators' Twitter Use

matt-koehler.com

Open at publisher

Quoted on this wiki

Every place a page here uses this source, in the order the words come in it.

  1. Page 1

    we suggest practical, holistic metrics that can be employed to help identify spam. Although educators use a variety of social media to advance their own learning, Twitter has been particularly popular among both educators organizing their own professional learning and the researchers studying them (e.g., Carpenter and Krutka 2014; Rosenberg, Greenhalgh et al. 2016; Tucker 2019; Veletsianos 2017). One of the features of Twitter that makes this platform especially helpful for creating professional learning commu- nities is the hashtag (i.e., a word or phrase, preceded by a hash or pound sign, that organizes content by topics). Education- focused Twitter hashtags provide educators with spaces to connect, share ideas and discuss education topics (Carpenter et al. 2018; Rosenberg et al. 2016). These spaces have been evaluated in terms of communities of practice (e.g., Britt and Paulus 2016; Gao and Li 2017), professional learning communities (e.g., Goodyear et al. 2019) and affinity spaces (e.g., Staudt Willet 2019; Rosenberg et al. 2016). Although there are important distinctions between these frameworks, the literature generally concludes that professional learning spaces created through Twitter hashtags can be of real value to educators. It is therefore unsurprising that some teacher educators now explicitly introduce Twitter and Twitter hashtags to preservice teachers (e.g., Carpenter 2015; Carpenter and Morrison 2018; Hsieh 2017; Luo et al. 2017). Yet, the same Twitter features that allow for the easy creation of spaces for professional learning also allow spam—“unsolic- ited, repeated actions that negatively impact other people” * Jeffrey P. Carpenter [email protected] K. Bret Staudt Willet [email protected] Matthew J. Koehler [email protected] Spencer P.

    In Community participation and contribution

  2. Page 1

    such spaces can also attract unwelcome Twitter traffic that complicates researchers’ attempts to explore and understand educators’ professional social media experiences. Although educators use a variety of social media to advance their own learning, Twitter has been particularly popular among both educators organizing their own professional learning and the researchers studying them (e.g., Carpenter and Krutka 2014; Rosenberg, Greenhalgh et al. 2016; Tucker 2019; Veletsianos 2017). One of the features of Twitter that makes this platform especially helpful for creating professional learning commu- nities is the hashtag (i.e., a word or phrase, preceded by a hash or pound sign, that organizes content by topics). Education- focused Twitter hashtags provide educators with spaces to connect, share ideas and discuss education topics (Carpenter et al. 2018; Rosenberg et al. 2016). These spaces have been evaluated in terms of communities of practice (e.g., Britt and Paulus 2016; Gao and Li 2017), professional learning communities (e.g., Goodyear et al. 2019) and affinity spaces (e.g., Staudt Willet 2019; Rosenberg et al. 2016). Although there are important distinctions between these frameworks, the literature generally concludes that professional learning spaces created through Twitter hashtags can be of real value to educators. It is therefore unsurprising that some teacher educators now explicitly introduce Twitter and Twitter hashtags to preservice teachers (e.g., Carpenter 2015; Carpenter and Morrison 2018; Hsieh 2017; Luo et al. 2017). Yet, the same Twitter features that allow for the easy creation of spaces for professional learning also allow spam—“unsolic- ited, repeated actions that negatively impact other people” * Jeffrey P. Carpenter [email protected] K. Bret Staudt Willet [email protected] Matthew J. Koehler [email protected] Spencer P.

    In Community participation and contribution

  3. Page 2

    Greenhalgh [email protected] 1 Elon University, 336 278 5969, Campus Box 2105, Elon, NC 27278, USA 2 Michigan State University, 620 Farm Ln, East Lansing, MI 48824, USA 3 University of Kentucky, 320 Lucille Little Library Bldg., Lexington, KY 40506, USA TechTrends https://doi.org/10.1007/s11528-019-00466-3 this low threshold allows for irrelevant, off-topic, commercial, or malicious mes- sages to easily enter the space. In any of these cases, however, identify- ing spam and establishing a plan for what to do with it is a key step, albeit an often-overlooked one. Background In the following sections, we describe some basic characteris- tics of spam as well as how personal and contextual factors may influence an individual’s perceptions of those characteristics. In doing so, we draw on selected sources from the broader litera- ture on spam while concentrating on how spam and spam-like activity have been acknowledged in Twitter research by educa- tional technology scholars. We conclude with a brief summary of how researchers have responded to spam. Characteristics of Spam Spam has been described as “undesirable text, whether repet- itive, excessive, or interfering” (Brunton 2013, p. xxii). This broad definition is helpful for acknowledging the wide range of forms that spam can take, but attention to some of the specific ways text can be undesirable (especially within an education-focused community) is necessary for fully under- standing this phenomenon. These are described below and exemplified in the faux tweets in Fig. 1. & Unintentionality: The same hashtag may mean different things to different people, leading to two different groups accidentally occupying the same space. Similarly, mis- takes or typos may lead to someone posting to an educa- tional hashtag entirely by accident. For example, in Fig. 1a, the “#maet” hashtag is used by participants in a Master of Arts in Educational Technology program, but the functionally-identical “#mæt” hashtag refers to the Danish word for “full” (see Greenhalgh et al. 2016). & Irrelevance: Hashtags are open spaces, which lowers bar- riers to participation for both educators and those speaking on other, unrelated topics. & Commercial nature: Among the unrelated messages that may enter an educational hashtag are those promoting commercial products (whether educational or not).

    In Community participation and contribution

  4. Page 2

    Greenhalgh [email protected] 1 Elon University, 336 278 5969, Campus Box 2105, Elon, NC 27278, USA 2 Michigan State University, 620 Farm Ln, East Lansing, MI 48824, USA 3 University of Kentucky, 320 Lucille Little Library Bldg., Lexington, KY 40506, USA TechTrends https://doi.org/10.1007/s11528-019-00466-3 Because popular Twitter hashtags attract attention from a considerable audience, such hashtags can prove an enticing target for spammers who attempt to redirect users’ attention. In any of these cases, however, identify- ing spam and establishing a plan for what to do with it is a key step, albeit an often-overlooked one. Background In the following sections, we describe some basic characteris- tics of spam as well as how personal and contextual factors may influence an individual’s perceptions of those characteristics. In doing so, we draw on selected sources from the broader litera- ture on spam while concentrating on how spam and spam-like activity have been acknowledged in Twitter research by educa- tional technology scholars. We conclude with a brief summary of how researchers have responded to spam. Characteristics of Spam Spam has been described as “undesirable text, whether repet- itive, excessive, or interfering” (Brunton 2013, p. xxii). This broad definition is helpful for acknowledging the wide range of forms that spam can take, but attention to some of the specific ways text can be undesirable (especially within an education-focused community) is necessary for fully under- standing this phenomenon. These are described below and exemplified in the faux tweets in Fig. 1. & Unintentionality: The same hashtag may mean different things to different people, leading to two different groups accidentally occupying the same space. Similarly, mis- takes or typos may lead to someone posting to an educa- tional hashtag entirely by accident. For example, in Fig. 1a, the “#maet” hashtag is used by participants in a Master of Arts in Educational Technology program, but the functionally-identical “#mæt” hashtag refers to the Danish word for “full” (see Greenhalgh et al. 2016). & Irrelevance: Hashtags are open spaces, which lowers bar- riers to participation for both educators and those speaking on other, unrelated topics. & Commercial nature: Among the unrelated messages that may enter an educational hashtag are those promoting commercial products (whether educational or not).

    In Community participation and contribution

  5. Page 2

    In any of these cases, however, identify- ing spam and establishing a plan for what to do with it is a key step, albeit an often-overlooked one. Background In the following sections, we describe some basic characteris- tics of spam as well as how personal and contextual factors may influence an individual’s perceptions of those characteristics. In doing so, we draw on selected sources from the broader litera- ture on spam while concentrating on how spam and spam-like activity have been acknowledged in Twitter research by educa- tional technology scholars. We conclude with a brief summary of how researchers have responded to spam. Characteristics of Spam Spam has been described as “undesirable text, whether repet- itive, excessive, or interfering” (Brunton 2013, p. xxii). This broad definition is helpful for acknowledging the wide range of forms that spam can take, but attention to some of the specific ways text can be undesirable (especially within an education-focused community) is necessary for fully under- standing this phenomenon. These are described below and exemplified in the faux tweets in Fig. 1. & Unintentionality: The same hashtag may mean different things to different people, leading to two different groups accidentally occupying the same space. Similarly, mis- takes or typos may lead to someone posting to an educa- tional hashtag entirely by accident. For example, in Fig. 1a, the “#maet” hashtag is used by participants in a Master of Arts in Educational Technology program, but the functionally-identical “#mæt” hashtag refers to the Danish word for “full” (see Greenhalgh et al. 2016). & Irrelevance: Hashtags are open spaces, which lowers bar- riers to participation for both educators and those speaking on other, unrelated topics. & Commercial nature: Among the unrelated messages that may enter an educational hashtag are those promoting commercial products (whether educational or not). they can—and often do—co-exist within a single message. communities have an explicit focus on certain products (e.g., #TLAP for Burgess’s 2012 book, Teach Like A Pirate, or #gafe for Google Apps for Education), what would typically be considered commercial spam may, in fact, be welcomed. In other scenarios, however, individual cases may be eval- uated differently by individual participants. A tweet using several hashtags to promote an educational workshop (see Greenhalgh 2018) may be of genuine interest to some teachers but seen by others as an unwelcome commercial intrusion. Indeed, a tweet could feature none of the characteristics de- scribed above but still be deemed “undesirable.” For example, many educational hashtags host regular, synchronous chats: hour-long conversations structured around pre-determined prompts (Evans 2015; Gao and Li 2017). Research has report- ed that some participants feel overwhelmed during high- volume chats (Britt and Paulus 2016; Luo et al. 2017); even though this may not correspond with one’s intuitive under- standing of spam, the sheer volume may still discourage par- ticipation just as offensive or commercial content would. Researcher Responses to Spam Researchers may also differ in their perceptions of spam— more importantly, they face the task of consistently identifying and intentionally responding to spam. As the bulk of this paper is dedicated to our suggestions for identifying spam, we focus here on the different ways that researchers may re- spond to spam. In some instances, spam itself may even be of interest. For instance, Twitter has been plagued by cyberviolence and misogyny (Nagle 2018), and researchers may be interested in documenting the degree to which such offensive or harassing spam is present in ostensibly education- focused Twitter spaces. Furthermore, the phenomenon of teacherpreneurship (e.g., Shelton and Archambault 2018) may lead researchers to investigate commercial messages originating with teachers.

    In Content-based spam filtering

  6. Page 5

    Researchers could consider raw number of hashtags, the percentage of tweets that contain more than one hashtag (i.e., a hashtag in addition to the one defining the space under investigation), or the average number of hashtags per tweet. & Level of hyperlinking: Spammers often include hyperlinks in their tweets in an attempt to drive traffic to certain websites (e.g., Lin and Huang 2013). For instance, a tweet may advertise goods for sale and include a hyperlink to the website where those goods could be purchased. Researchers can therefore analyze the raw number of links, the average number of links per tweet, or the per- centage of tweets that include links. & Bot-like activity: Automated accounts (i.e., bots) can be the sources of a lot of the commercial spam found in educator social media spaces. Researchers and Twitter, Inc. continuously try to develop tools and techniques to detect bots (Chen et al. 2016), while many bot creators simultaneously change techniques to thwart detection. Tools such as Botometer (https://botometer.iuni.iu.edu/) examine a user’s Twitter activity and estimate the probability that the user is a bot. These metrics do not align directly with the types and char- acteristics of spam that we outlined earlier. Characteristics such as offensiveness and irrelevance are often judgement calls that require interpretation and context not easily detect- able in large Twitter datasets. Furthermore, as previously de- scribed, personal and contextual factors play a key role in the difference between potential spam and actual spam. Nonetheless, irrelevant, commercial, offensive, and impersonal forms of spam frequently leave behind digital traces that are reflected in the above metrics. Thus, these TechTrends exclusive reliance upon one quantitative metric (or even several) would likely result in failures to identify some spam while also falsely identifying some accounts as spammers. These decisions should align with the research questions under in- vestigation, the nature of the data itself and the level of cer- tainty in the decision being made. For instance, research ques- tions related to the full range of experiences, uses and pur- poses of an education-focused hashtag are likely to exclude very little data from analysis. After all, some spam can unfor- tunately be part of educators’ typical Twitter experiences. However, if research questions are directed towards the pro- fessional interactions that Twitter can facilitate, spammers who fail to actually engage with educators could reasonably be excluded from the dataset. In a previous conference paper (Carpenter et al. 2019), we described the use of this holistic approach to identify spam in one dataset. In this section, we describe different practical benefits of identifying this spam and the associated consider- ations researchers can make. In doing so, we expand our de- scription of that instance of spam identification and introduce two other implementations of our proposed approach. We present these examples of identifying spam and making deci- sions about its inclusion not to suggest a one-size-fits-all ap- proach for educational researchers, but instead as an opportu- nity for scholars to consider appropriate metrics and strategies related to the spam that they may encounter in their own re- search. A holistic decision-making process can lead re- searchers in different directions, depending on a study’s focus and research questions. Practical Use #1: Allowing for Accurate Comparisons In some studies, being able to accurately identify and remove spam can be essential for making valid comparisons between communities. Carpenter and colleagues (2018) aimed to de- pict the landscape of education-focused Twitter hashtags by comparing and contrasting the traffic from a set of 16 hashtags over a 13-month timespan.

    In Email risk classification

  7. Page 7

    Similarly, in this particular case, a holistic evaluation helped in- crease confidence that key measures would not be unduly swayed by a few users’ activity. We recommend that researchers plan to implement spam detection at the outset of data analysis. In this way, spam metrics can serve as checks for participation outliers and determine how much the prolific contributors affect the overall description of activity (whether or not they are deter- mined to be spammers). Sometimes, as in the cases of #satchat or #BFC530, participation inequality greatly impacts findings; in other cases, such as for #Edchat, the contributions of the 10 most prolific contributors were on topic and did not dramatically im- pact the description of the overall hashtag traffic. Practical Use #3: Focusing on Spam-like Behavior Unlike the previous two cases, we describe here how our prac- tical metrics can identify potential spammers not in order to remove them but rather to focus on them as an actual phenom- enon. Throughout these cases, we have focused on how these spam metrics can be helpful in identifying users whose Twitter activity is far outside the norm. Regardless of whether this extreme activity is ultimately considered to be spam, researchers may find value in studying these outliers. For example, Staudt Willet (2019) suggested a need for future research to study self- promotional behavior on Twitter, which may appear to some to be spam. We therefore revisit the #Edchat data described above with the goal of identifying self-promoting users. Given this goal, our focus is not just on volume metrics but also on interaction metrics such as the average likes, retweets and replies per tweet. these metrics can also help identify specific cases for further study. Social media such as Twitter may also improve their own policing of certain types of spam. In the near future, however, we see value for many research contexts in the utilization of a combination of metrics and a final holistic hu- man decision to find spam and consider its removal. Users who access education-focused Twitter hashtags have diverse pur- poses and motivations (e.g., Carpenter and Krutka 2014), and the multiple flavors of spam that can affect those spaces can confound purely algorithmic methods. TechTrends

    In Metrics