...trovi il tuo blog linkato nelle dispense di due università americane (Central Florida e Central Connecticut). Nice!
Il mio post sul workaround per far funzionare Dev-CPP sotto Vista (dio maledica Microsoft per le sue trovate balzane) è diventato un po' il fix de facto in giro per internet, e ha portato più traffico e commenti delle altre centinaia di post che ho realizzato, anche quelle probabilmente più utili come l'implementazione di Aho-Corasick in PHP (la prima open-source, afaik).
Quest'ultima è un fenomeno interessante, in quanto mi ha provocato solamente un'ondata di spam pornografico nei commenti... in giapponese. Qualcuno deve aver linkato il post su un sito del sol levante (che non riesco a recuperare), e da lì è partito un assalto che, dopo quasi un anno, non accenna ad arrestarsi.
Showing posts with label university. Show all posts
Showing posts with label university. Show all posts
Tuesday, 8 February 2011
Wednesday, 2 February 2011
UNIMIB.IT on Webometrics: new strong performance
Webometrics publishes, in a biannual fashion, a ranking of universities' websites, produced by a formula that takes into account the size of the site, the number of downloadable documents, the visibility online and the number of google scholar entries. We still have a positive trend and during the latter half we went up 145 positions!
Once again Milano-Bicocca pays for the size of the site, intended as the "number of pages recovered on four different search engines". We have a lot of information spread in too many satellite websites, like faculties and departments.
A very positive score comes from Google Scholar indicator: Bicocca-related research documents have grown a lot during the latter years: in 2007 we were in 682th position, now we have reached the 277th. It's a very good result, probably due to the introduction of BOA service, despite the fact that Milano-Bicocca is one of the youngest universities.
Since this january, Milano-Bicocca is featured in the Top-500 ranking, placing in the 477th position and widening Italy's slice a little bit.
![]() |
| Milano-Bicocca's ranking during the years; the lower, the better |
Once again Milano-Bicocca pays for the size of the site, intended as the "number of pages recovered on four different search engines". We have a lot of information spread in too many satellite websites, like faculties and departments.
A very positive score comes from Google Scholar indicator: Bicocca-related research documents have grown a lot during the latter years: in 2007 we were in 682th position, now we have reached the 277th. It's a very good result, probably due to the introduction of BOA service, despite the fact that Milano-Bicocca is one of the youngest universities.
Since this january, Milano-Bicocca is featured in the Top-500 ranking, placing in the 477th position and widening Italy's slice a little bit.
Friday, 23 July 2010
Milano-Bicocca improves its Webometrics ranking
Webometrics ranking of world universities shows information about world universities, ranked according to indicators measuring web presence, link visibility, size and number of documents available for download. Since 2004 this chart is done by the Spanish National Research Council, and gets updated every six months.
Milan-Bicocca's website has a strongly positive trend as we raised 325 positions since 2007; during the last year we jumped 19 positions.
Observing the data, one notes that we pay for the size of our site, intended as the "number of pages recovered on four different search engines". We have a lot of information spread on satellite website, like faculties. It's not clear, according to the methodological note, if those pages are counted or not.
Moreover, there's an important issue connected with our research institutions. Isidro Aguillo, editor of the ranking, writes: "the rankings of universities of Germany, France, Italy and Spain come later [than British and nordic ones], mostly due to the fact that research in this countries is developed by independent institutions (Max Planck, CNRS, CNR, CSIC)". Therefore, our ranking is strongly lowered by this kind of structure.
Monday, 7 June 2010
Mergifier: a bayesian RSS filter
The RSS feed is a popular format used to publish information. XML based, therefore machine-readable, it was thought as a "container" of news from frequently updated web pages. Personally, I love using iGoogle as aggregator to keep, in a single screen, lot of feeds that keep me constantly updated about intersting facts happening around the world.
We introduced early this technology on Bicocca's website, to spread interesting facts in an aggregate way; we placed RSS feeds in pages featuring news, events and so forth.
We recently talked about this system and we did not find it optimal. We think that it could be more effective to route all this informations into single feeds intended to particular audience, instead of fragmenting everything into many rarely-changed small feeds. Stefano has partially mitigated the problem by creating a feed's map, but we agreed that is sub-optimal: we still prefer something able to gather everything in a single channel and "skim" the news according to the reader's interest.
A similar project is Yahoo! Pipes, that I discovered during the course of knowledge representation. YP allows, among other, to "sum up" contribution of different feeds into a single one but cannot classify the outgoing items (or, at least, I can't find anything like that :) ). Reading some post from the project's board, it seems I'm not the only one lamenting this lack. Not a problem, because in the meanwhile I rediscovered the wheel.
"Mergifier", that's the name I gave to the software, draws from an arbitrary number of RSS feeds (and ATOM as well) and integrates them in a big cauldron. The administrator - that is us, at the moment - is allowed to create as many categories as he wants and every item can be classified in one or more of those. Naturally, I'm not talking about a trivial table, because we don't want to classify every entry by hand. That's why I used a bayesian filter.
More ham for everyone
The idea comes from the bayesian spam filtering used in software like Mozilla Thunderbird: whenever you receive an unwanted letter, you classify it as "junk". The client stores words' frequencies and use those informations to determine the probability of future e-mail's to be spam.
The probability estimation is quite simple:

where p1...Pn is the individual probability of a word to be present in a "junk e-mail". How can we determine that? By using the following formula:
where
This method works stunningly fine; we're testing the results in order to release an alpha during the next weeks :) .
We introduced early this technology on Bicocca's website, to spread interesting facts in an aggregate way; we placed RSS feeds in pages featuring news, events and so forth.
We recently talked about this system and we did not find it optimal. We think that it could be more effective to route all this informations into single feeds intended to particular audience, instead of fragmenting everything into many rarely-changed small feeds. Stefano has partially mitigated the problem by creating a feed's map, but we agreed that is sub-optimal: we still prefer something able to gather everything in a single channel and "skim" the news according to the reader's interest.
A similar project is Yahoo! Pipes, that I discovered during the course of knowledge representation. YP allows, among other, to "sum up" contribution of different feeds into a single one but cannot classify the outgoing items (or, at least, I can't find anything like that :) ). Reading some post from the project's board, it seems I'm not the only one lamenting this lack. Not a problem, because in the meanwhile I rediscovered the wheel.
"Mergifier", that's the name I gave to the software, draws from an arbitrary number of RSS feeds (and ATOM as well) and integrates them in a big cauldron. The administrator - that is us, at the moment - is allowed to create as many categories as he wants and every item can be classified in one or more of those. Naturally, I'm not talking about a trivial table, because we don't want to classify every entry by hand. That's why I used a bayesian filter.
More ham for everyone
The idea comes from the bayesian spam filtering used in software like Mozilla Thunderbird: whenever you receive an unwanted letter, you classify it as "junk". The client stores words' frequencies and use those informations to determine the probability of future e-mail's to be spam.
The probability estimation is quite simple:
where p1...Pn is the individual probability of a word to be present in a "junk e-mail". How can we determine that? By using the following formula:
- Pr(S|W) is the probability of e-mail to be spam
- Pr(W|S) is the chance of finding the word in a junk e-mail
- Pr(S) is the "a priori" probability of a mail to be spam
- Pr(W|H) is the chance of finding the word in a good email
- Pr(H) is the "a priori" probability of a mail to be ham (that is the way "not spam" e-mail are called!)
This method works stunningly fine; we're testing the results in order to release an alpha during the next weeks :) .
Saturday, 17 April 2010
Un percettrone per il portalone
Da un po' di tempo a questa parte mi frulla l'idea di rendere in qualche modo "intelligente" le generazione del primo piano del portale www.unimib.it.
Fresco del corso di machine learning del prof. Mauri ho deciso di tentare l'implementazione di un software in grado di imparare dagli esempi - storici e futuri - la mentalità con cui realizziamo giornalmente il primo piano, arrivando un giorno a realizzarlo in maniera automatica. A.N.N.N.E, questo il nome del programma, attualmente sfrutta una semplice rete neurale feed-forward. Nell'esperimento cui si riferisce lo screenshot ha avuto in pasto circa 1700 esempi, tutti classificati in base ad una manciata di attributi.
Gli esempi sono stati sufficienti per avere circa il 22% di errori di classificazione; non è male ma temo che il problema, più che della rete in sè, sia nella limitatezza del modello che non tiene in alcun modo conto della componente temporale.
Può capitare infatti che qualche notizia/evento importante non finisca nel primo piano per banali ragioni di congestionamento. Inoltre, gli eventi è più probabile che finiscano in primo piano nei giorni immediatamente precedenti l'evento stesso. Per le notizie vale l'opposto: più tempo è trascorso e più è probabile la discesa negli avvisi normali. Entrambi gli aspetti non sono considerati al momento, e sicuramente influiscono negativamente nel risultato.
L'inizio è comunque stimolante.
Monday, 4 May 2009
...and the story ends
I think this is the time to assess a balance of these years.
I'm grown, surely. Grown as a man and programmer. The stage (and thesis) I chose were challenging and helped me understanding who I am and what I could do. I learnt lot from this experience, and from all these years. I'm surely a better computer scientist than before.
Talking about our professors, it turned out that many of our teachers are extraordinary people, and a few are extraordinary assholes. Shit happens, I suppose it's the same eveywhere. The coolest course was, of course, cybernetics; the worst.. mm.. I suppose databases. In the middle, some wonderful, some useless, some I'll never forget, and a couple that changed my way of thinking forever. There were a lot of mind-expanding matters and a few simply mind-blowing.
It's nice that, even in my small thesis, a little piece of Gauss' work appeared: my last convolution filter used a discrete gaussian kernel in order to obtain a better smoothing! :)
Thursday, 26 February 2009
Wednesday, 16 January 2008
Signals of hope
In the meanwhile, students at Università la Sapienza prevented Pope's visit in the name of Science.
Good news is that there's still hope in this old, tired and stinky country.
Bad news is that our politicians raised the usual, transversal choir of disapproval. Even our prime minister Prodi defends such a criminal that stomped on justice even to help Mafia.
Subscribe to:
Posts (Atom)




