Showing posts with label university. Show all posts
Showing posts with label university. Show all posts

Tuesday, 8 February 2011

Realizzi di esser stato d'aiuto quando...

...trovi il tuo blog linkato nelle dispense di due università americane (Central Florida e Central Connecticut). Nice!

Il mio post sul workaround per far funzionare Dev-CPP sotto Vista (dio maledica Microsoft per le sue trovate balzane) è diventato un po' il fix de facto in giro per internet, e ha portato più traffico e commenti delle altre centinaia di post che ho realizzato, anche quelle probabilmente più utili come l'implementazione di Aho-Corasick in PHP (la prima open-source, afaik).

Quest'ultima è un fenomeno interessante, in quanto mi ha provocato solamente un'ondata di spam pornografico nei commenti... in giapponese. Qualcuno deve aver linkato il post su un sito del sol levante (che non riesco a recuperare), e da lì è partito un assalto che, dopo quasi un anno, non accenna ad arrestarsi.

Wednesday, 2 February 2011

UNIMIB.IT on Webometrics: new strong performance

Webometrics publishes, in a biannual fashion, a ranking of universities' websites, produced by a formula that takes into account the size of the site, the number of downloadable documents, the visibility online and the number of google scholar entries. We still have a positive trend and during the latter half we went up 145 positions!

Milano-Bicocca's ranking during the years; the lower, the better


Once again Milano-Bicocca pays for the size of the site, intended as the "number of pages recovered on four different search engines". We have a lot of information spread in too many satellite websites, like faculties and departments.

A very positive score comes from Google Scholar indicator: Bicocca-related research documents have grown a lot during the latter years: in 2007 we were in 682th position, now we have reached the 277th. It's a very good result, probably due to the introduction of BOA service, despite the fact that Milano-Bicocca is one of the youngest universities.

Since this january, Milano-Bicocca is featured in the Top-500 ranking, placing in the 477th position and widening Italy's slice a little bit.

Friday, 23 July 2010

Milano-Bicocca improves its Webometrics ranking



Webometrics ranking of world universities shows information about world universities, ranked according to indicators measuring web presence, link visibility, size and number of documents available for download. Since 2004 this chart is done by the Spanish National Research Council, and gets updated every six months.

Milan-Bicocca's website has a strongly positive trend as we raised 325 positions since 2007; during the last year we jumped 19 positions.

Observing the data, one notes that we pay for the size of our site, intended as the "number of pages recovered on four different search engines". We have a lot of information spread on satellite website, like faculties. It's not clear, according to the methodological note, if those pages are counted or not.

Moreover, there's an important issue connected with our research institutions. Isidro Aguillo, editor of the ranking, writes: "the rankings of universities of Germany, France, Italy and Spain come later [than British and nordic ones], mostly due to the fact that research in this countries is developed by independent institutions (Max Planck, CNRS, CNR, CSIC)". Therefore, our ranking is strongly lowered by this kind of structure.

Monday, 7 June 2010

Mergifier: a bayesian RSS filter

The RSS feed is a popular format used to publish information. XML based, therefore machine-readable, it was thought as a "container" of news from frequently updated web pages. Personally, I love using iGoogle as aggregator to keep, in a single screen, lot of feeds that keep me constantly updated about intersting facts happening around the world.

We introduced early this technology on Bicocca's website, to spread interesting facts in an aggregate way; we placed RSS feeds in pages featuring news, events and so forth.

We recently talked about this system and we did not find it optimal. We think that it could be more effective to route all this informations into single feeds intended to particular audience, instead of fragmenting everything into many rarely-changed small feeds. Stefano has partially mitigated the problem by creating a feed's map, but we agreed that is sub-optimal: we still prefer something able to gather everything in a single channel and "skim" the news according to the reader's interest.

A similar project is Yahoo! Pipes, that I discovered during the course of knowledge representation. YP allows, among other, to "sum up" contribution of different feeds into a single one but cannot classify the outgoing items (or, at least, I can't find anything like that :) ). Reading some post from the project's board, it seems I'm not the only one lamenting this lack. Not a problem, because in the meanwhile I rediscovered the wheel.

"Mergifier", that's the name I gave to the software, draws from an arbitrary number of RSS feeds (and ATOM  as well) and integrates them in a big cauldron. The administrator - that is us, at the moment - is allowed to create as many categories as he wants and every item can be classified in one or more of those. Naturally, I'm not talking about a trivial table, because we don't want to classify every entry by hand. That's why I used a bayesian filter.

More ham for everyone

The idea comes from the bayesian spam filtering used in software like Mozilla Thunderbird: whenever you receive an unwanted letter, you classify it as "junk". The client stores words' frequencies and use those informations to determine the probability of future e-mail's to be spam.

The probability estimation is quite simple: 


where p1...Pn  is the individual probability of a word to be present in a "junk e-mail". How can we determine that? By using the following formula:

where
  • Pr(S|W) is the probability of e-mail to be spam
  • Pr(W|S) is the chance of finding the word in a junk e-mail
  • Pr(S) is the "a priori" probability of a mail to be spam
  • Pr(W|H) is the chance of finding the word in a good email
  • Pr(H) is the "a priori" probability of a mail to be ham (that is the way "not spam" e-mail are called!)
The idea was to replace the concept of spam with our own categories, and compute those probabilities "online" by fetching from the feeds and letting us classify them. Naturally, the most common words are blacklisted in order to lower the work of database.

This method works stunningly fine; we're testing the results in order to release an alpha during the next weeks :) .

Saturday, 17 April 2010

Un percettrone per il portalone


Da un po' di tempo a questa parte mi frulla l'idea di rendere in qualche modo "intelligente" le generazione del primo piano del portale www.unimib.it.

Fresco del corso di machine learning del prof. Mauri ho deciso di tentare l'implementazione di un software in grado di imparare dagli esempi - storici e futuri - la mentalità con cui realizziamo giornalmente il primo piano, arrivando un giorno a realizzarlo in maniera automatica. A.N.N.N.E, questo il nome del programma, attualmente sfrutta una semplice rete neurale feed-forward. Nell'esperimento cui si riferisce lo screenshot ha avuto in pasto circa 1700 esempi, tutti classificati in base ad una manciata di attributi.

Gli esempi sono stati sufficienti per avere circa il 22% di errori di classificazione; non è male ma temo che il problema, più che della rete in sè, sia nella limitatezza del modello che non tiene in alcun modo conto della componente temporale.

Può capitare infatti che qualche notizia/evento importante non finisca nel primo piano per banali ragioni di congestionamento. Inoltre, gli eventi è più probabile che finiscano in primo piano nei giorni immediatamente precedenti l'evento stesso. Per le notizie vale l'opposto: più tempo è trascorso e più è probabile la discesa negli avvisi normali. Entrambi gli aspetti non sono considerati al momento, e sicuramente influiscono negativamente nel risultato.

L'inizio è comunque stimolante.

Monday, 4 May 2009

...and the story ends

56 months have passed since my second enrollment, but this time it worked! I'm finally graduated in computer science with a final score of 99/110.

I think this is the time to assess a balance of these years.

I'm grown, surely. Grown as a man and programmer. The stage (and thesis) I chose were challenging and helped me understanding who I am and what I could do. I learnt lot from this experience, and from all these years. I'm surely a better computer scientist than before.

Talking about our professors, it turned out that many of our teachers are extraordinary people, and a few are extraordinary assholes. Shit happens, I suppose it's the same eveywhere. The coolest course was, of course, cybernetics; the worst.. mm.. I suppose databases. In the middle, some wonderful, some useless, some I'll never forget, and a couple that changed my way of thinking forever. There were a lot of mind-expanding matters and a few simply mind-blowing.

Graduation let me also know lot of special people: living and among.. the dead! Now I can recognize the importance, for instance, of people like Gauss, whose name flows constantly along our graduation path: from algorythms to cybernetics, through linear algebra, probability and physics, it seems that everythings sprung out of his incredible mind.

It's nice that, even in my small thesis, a little piece of Gauss' work appeared: my last convolution filter used a discrete gaussian kernel in order to obtain a better smoothing! :)

Thursday, 26 February 2009

俳句

Race is over:
a long breath.
Shivers in the sun.

Wednesday, 16 January 2008

Signals of hope

Italian worst minister ever, Clemente Mastella, presented resignation this morning after the nth scandal involving him. After months of sufference and bad administration our Justice Ministry will be released.

In the meanwhile, students at Università la Sapienza prevented Pope's visit in the name of Science.

Good news is that there's still hope in this old, tired and stinky country.

Bad news is that our politicians raised the usual, transversal choir of disapproval. Even our prime minister Prodi defends such a criminal that stomped on justice even to help Mafia.