• No results found

Analysis of estimated parameters

5.5 Experimental results

5.5.4 Analysis of estimated parameters

This subsection describes the analysis of the estimated parameters of the CRD models as well as its external influence functions.

Circadian Rhythm and Aging. As a by-product of modelling each petition using the Circadian with Rise and Decay (CRD) model given in Eq. 1, we obtain a distribution for each parameter across all petitions. These distributions are shown in Figure 5.12, where we are separating failed petitions from successful ones, as well as a special case of successful petitions, which are the ones promoted on the front page.

As expected, we observe that the intensity parametera, which corresponds to the vertical shift of the series of signatures per unit of time, is higher for successful petitions that for unsuccessful ones. Interestingly, the amplitude parameterbshows that the oscillations of the series are larger for failed petitions. The growth parameterk, which influences the day at which a petition reaches its peak, shows that successful petitions tend to be more popular early on in comparison with failed petitions, and that the peak of the petitions that are promoted on the front page happens later in time—usually at the moment the petition ranks the highest on the front page. The decay parameterτcan be much larger for successful petitions, meaning that they sustain interest for a longer period of time (in the model this appears ase−t/τ). Finally, most of the petitions have a

Chapter 5. Predicting the Success of Online Petitions 0.0 0.2 0.4 0.6 0.8 1.0 F raction of F ailed P etitions Median 0.0 0.2 0.4 0.6 0.8 1.0 F raction of Successful P etitions 0 20 40 60 80 100 Intensity a 0.0 0.2 0.4 0.6 0.8 1.0 F raction of F ront P age Promoted 0 5 10 15 20 25 Amplitude b 0 1 2 3 4 5 6 7 Growth k 0 10203040506070 Decay τ 0 2 4 6 8 10 12 Phase φ

Figure 5.12 – Distributions of parameter estimations for failed petitions (top) and successful petitions (middle). We also consider separately petitions promoted on the front page, all of them successful (bottom). Details about each parameter are provided in Section 5.4.1.

in the USA and signed by people in the same country, in time zones that are close to each other (the distributions are almost equal so they are omitted from the figure).

Self-Excitation vs External Influence. Our model uses a time window of sizeTmemhours, which allows to incorporate information from the recent past in its estimation of the future. Each of the coefficients for the influence of self-excitationcself(i ), social mediacsm(i ), and front-page effectcfront(i )can be seen as a time-indexed vector reflecting the importance of different moments of the recent past for each specific influence across successful petitions. If we are predicting the popularity ont+ 1hour, the influence function corresponds to the vector of sizeTmemthat contains the effect of eacht−i hour of observation from self-influence, social media or front-page influence, wherei= 0,1,...,Tmem. The centroids of these vectors are shown in Figure 5.13.

0 2 4 6 8 10 12 −20 2 4 6

Influence

absolute

Self- c

self

(i)

Failed Successful

0 2 4 6 8 10 12

Hours i before prediction time t+ 1

0 6 12 18

24

Soc. media c

sm

(i)

Failed Successful 0 2 4 6 8 10 0 42 84

126

Front-page c

front

(i)

Successful

Figure 5.13 – Influence function estimation for self-excitation, the influence of social media, and the

front-page effect. A value ofi on the X axis refers to the median influence of this aspecti hours in the

past. On each plot, the Y axis presents an absolute scale for successful and failed petitions and shows the multiplicative effect on the number of signatures. Petitions promoted on the home page are all successful. Several interesting observations can be made from Figure 5.13. First, self-excitation seems to be

5.6. Conclusions

largely memory-less, with the immediately preceding step being the most influential element. Second, social media (Twitter in this case) has an influence that can last up to four hours for the successful petitions, and peaks at about 3 hours; this means a posting at timet mostly affects the signature rate between timest+2h andt+3h. Failed petitions seems to be less affected by social

media with a short memory of 1 hour. Third, the front-page effect has an effect that lasts about two hours. In absolute terms, social media has a stronger effect than self excitation, and being featured on the front page has a stronger effect than social media activity: it results in a big boost of signatures, consistently with our observations from Section 5.3.5.

5.6

Conclusions

Online user engagement is a complex phenomenon, which can be better captured when consider- ing potential influences that might be affecting it. In general, interdependent phenomena across websites are less studied than phenomena happening on a particular website. In this chapter, we studied an important form of engagement—signing an online petition—we modeled two external influences: activity on social media, and promotion to front page. In both cases, we demonstrated significant improvement in modelling and predicting engagement when those influences are taken into account. In addition, we showed that the circadian rhythm of human activity, and the fact that interest decays over time, also need to be considered.

We analyzed the effect of social media and found it to be impactful in two ways. First, at a micro level, as demonstrated by the matching of people signing a petition and then posting about it shortly afterwards. Second, at a macro level, where we analyzed the effect of Twitter on the signature rate using a Granger causality test, and showed significant improvement in prediction accuracy when using social media—improvements that are particularly important to reduce the amount of time/data needed to perform an accurate prediction. We were also able to determine that the effect of an increase in postings on Twitter lasts for about 5-6 hours and peaks at about 3-4 hours. These findings are probably relevant beyond online petitions, as many campaigners in social media (e.g., advocating for brands, causes, or candidates) also perform similar activities in order to boost user engagement.

Specifically for online petitions, we showed that online petitions that are successful tend to peak early and to continue receiving attention for longer. In other words, it is not just about having a “strong start,” but about being able to sustain this engagement day after day. Petitions can be boosted by activity on social media, and/or by featuring them prominently to a large audience of potential signatories, as demonstrated by the front page effect that we have modeled and measured. These findings are relevant for people running other types of campaigns, and may be particularly important for crowdfunding campaigns.

In general, running a successful campaign on the web requires sustained attention and punctual interventions. In that context, interpretable models that can provide actionable insight about how

Chapter 5. Predicting the Success of Online Petitions

provide small advantages in terms of prediction accuracy.

Recommendation for the future campaigners. First, we have seen that most of the successful

petitions experience increased user participation during the first initial days, thus, it is important to prepare the material and the meticulous plan on how to engage more people and explore the possibilities to using multiple channels to convey the idea of the petitions as early as possible.

Second, we recommend the activists to establish the connection with the petition platforms’

owners and request them to feature the petition on the front page. This showed to have the strongest effect on the user gain compared to social media. Third, in this work we have not made a great distinction between various topics of the petitions, however, we have seen some evidence that users are less likely to post sensitive topics (LGBTQA, women rights, abuse) on social media. Therefore, other means to promote and spread the information shall be found for such topics.

Future Work. We believe that this chapter is an important step towards better modelling and predicting how reinforced information spreads online. It can be extended in a number of ways. In terms of new methods, it would be interesting to explore how the effects of several petitions on each other could be modeled, and how social media communities, defined both topically and through network structures, could be incorporated into our models. Moreover, impact functions could be represented through parametric distribution functions. In terms of enhancing the prediction accuracy, further sources of social media, and new features, could easily be incorporated into our model. Since we are modelling the petitions at an individual level, it might also be interesting to build and compare our model to a batch model and apply it over specific clusters of petitions. Finally, a prediction using a stochastic Hawkes process might be compared to the deterministic one presented in this chapter.

Part

6

Efficient Document Filtering Using

Vector Space Topic Expansion and

Pattern-Mining

The Case of Event Detection in Microposts

Automatically extracting information from social media is challenging given that social content is often noisy, ambiguous, and inconsistent. However, as many stories break on social channels first before being picked up by mainstream media, developing methods to better handle social content is of utmost importance. In this chapter, we propose a robust and effective approach to automatically identify microposts related to a specific topic defined by a small sample of reference documents. Our framework extracts clusters of semantically similar microposts that overlap with the reference documents, by extracting through frequent pattern mining combinations of key features that define those clusters. This allows us to construct compact and interpretable representations of the topic, dramatically decreasing the computational burden compared to clas- sical clustering and k-NN-based machine learning techniques and producing highly-competitive results even with small training sets (less than 1’000 training objects). Our method is efficient and scales gracefully with large sets of incoming microposts. We experimentally validate our approach on a large corpus of over 60M microposts, showing that it significantly outperforms state-of-the-art techniques.