These are the two things I am doing these days, apart from regular work at office.
1. Developing the next version of HarvestMan
2. Developing the next version of HarvestMan
Well, that is right... It was not a typing mistake :-)
HarvestMan is going to get a major update coming May, and it will be the result of more than 1.5 years of work. In fact the program has changed so much that I am changing the major version. It is going to be HarvestMan-2.0.
There are certain surprises with this HarvestMan release. Some of the interesting changes for NLP and Computational Linguistics programmers will be the addition of a plugin API that makes developing extensions for HarvestMan a breeze. In fact, the current CVS of HarvestMan already features an extension which binds it to an existing open source indexing engine. Apart from that, the program features a number of changes from the earlier version (1.4.6) so that it is almost on its way to becoming a "platform for web crawler software development", where I envision it to be.
Apart from that, this time HarvestMan will consist of two apps in one - that's right. There will be two applications using the same codebase. The crawler application (HarvestMan, of course) and a brand new web downloader application which supports multipart downloads. Let the name of this application be a mystery for the time being. The application just might change the command-line download experience of Unix/Linux users from the typical wget one. I will write more about it in the coming weeks.
In fact all of this is already in CVS. Anyone interested can checkout the latest code from berlios repository using anonymous CVS. There is not much documentation apart from the documentation in the code, but the code is pretty stable at the moment.
This should get released by mid May.
Wednesday, April 25, 2007
Tuesday, April 24, 2007
Open source and Innovation
What makes a successful open source project ? What makes a successful open source business ? How does successful open source projects make the transition to good business which still keep with the spirit of open source and open standards ?
These are questions any developer who is a serious contributor to open source would be interested in. Even if you are not an open source contributor, you will be interested in these questions, if your company is working in open source.
According to my experience so far, with my own projects and with some of the international projects I have participated in, a successful open source projects brings some new ideas to the table. New ideas need not be confused with new way of doing things or new I.P. It can be a new implementation of an existing protocol, it can be an open implementation of a proprietary standard, or it can be a project that uses existing open source components or applications to solve an existing problem (or a new one) in interesting and innovative ways. These need not always generate new kinds of intellectual property.
Yes, the key here is innovation. Good open source projects bring a fresh way of solving existing problems; they give a fresh perspective to existing way of doing things. Sometimes they are able to rewrite the rules by capturing the imagination of many hundreds of developers and thousands of supporting community members - a good example is the Firefox community. Some times, it will be a rather closeted group of skilled people in a rather niche area who finds a void in the experience of open source applications/operating systems and tries to fill the gap - a good example is the Beryl/Compiz projects which are working hard to bring display compositing to the Linux and open source crowd.
However, a common thread to all these project is this - they innovate. They innovate in fresh ideas, simplifying user experience and sometimes on performance. They often open up an entire new facet to an existing problem which makes programming a joy.
What do successful companies in open source have in common ? They understand the importance of keeping the developer crowd happy. They are keen to become good citizens of the open source community and contribute either their manpower or projects to the community - some do both. They understand that it is important not to just become consumers of open source but also stakeholders and participants.
When a company fails to understand this, or fails to create a working, effective developer policy towards open sourcing, it is prone to be assigned the category of a second or third rate citizen in the open source community. By just becoming a consumer of open source and not contributing enough, it risks alienating the coding crowd who tends to think of the company as a predator, not as an ally.
Most often, such companies never learn to use open source the right way too. By not participating enough, they fail to understand the driving force behind open source and people working in such projects. This in turn makes them less effective users of such software. For example, a company that brands itself as an open source integrator can never be quite effective if it does not understand the open source projects it is integrating and does not contribute developer resources to such projects; in fact, it is not even necessary to contribute directly most of the time. Indirect participation such as hosting meetings, contributing tools, toolchains and providing a platform for discussion and creation of new ideas are also good ways of contribution.
A company not doing any of these and still claiming to work in open source is somehow not doing the right thing. Such strategies are doomed to fail in the long term and even prove counter productive. In the long run a company like this is bound to move away from open source or bound to fail.
These are questions any developer who is a serious contributor to open source would be interested in. Even if you are not an open source contributor, you will be interested in these questions, if your company is working in open source.
According to my experience so far, with my own projects and with some of the international projects I have participated in, a successful open source projects brings some new ideas to the table. New ideas need not be confused with new way of doing things or new I.P. It can be a new implementation of an existing protocol, it can be an open implementation of a proprietary standard, or it can be a project that uses existing open source components or applications to solve an existing problem (or a new one) in interesting and innovative ways. These need not always generate new kinds of intellectual property.
Yes, the key here is innovation. Good open source projects bring a fresh way of solving existing problems; they give a fresh perspective to existing way of doing things. Sometimes they are able to rewrite the rules by capturing the imagination of many hundreds of developers and thousands of supporting community members - a good example is the Firefox community. Some times, it will be a rather closeted group of skilled people in a rather niche area who finds a void in the experience of open source applications/operating systems and tries to fill the gap - a good example is the Beryl/Compiz projects which are working hard to bring display compositing to the Linux and open source crowd.
However, a common thread to all these project is this - they innovate. They innovate in fresh ideas, simplifying user experience and sometimes on performance. They often open up an entire new facet to an existing problem which makes programming a joy.
What do successful companies in open source have in common ? They understand the importance of keeping the developer crowd happy. They are keen to become good citizens of the open source community and contribute either their manpower or projects to the community - some do both. They understand that it is important not to just become consumers of open source but also stakeholders and participants.
When a company fails to understand this, or fails to create a working, effective developer policy towards open sourcing, it is prone to be assigned the category of a second or third rate citizen in the open source community. By just becoming a consumer of open source and not contributing enough, it risks alienating the coding crowd who tends to think of the company as a predator, not as an ally.
Most often, such companies never learn to use open source the right way too. By not participating enough, they fail to understand the driving force behind open source and people working in such projects. This in turn makes them less effective users of such software. For example, a company that brands itself as an open source integrator can never be quite effective if it does not understand the open source projects it is integrating and does not contribute developer resources to such projects; in fact, it is not even necessary to contribute directly most of the time. Indirect participation such as hosting meetings, contributing tools, toolchains and providing a platform for discussion and creation of new ideas are also good ways of contribution.
A company not doing any of these and still claiming to work in open source is somehow not doing the right thing. Such strategies are doomed to fail in the long term and even prove counter productive. In the long run a company like this is bound to move away from open source or bound to fail.
Labels:
collaboration,
communities,
innovation,
open culture,
open source
Tuesday, July 04, 2006
Some ordered thoughts on Gnosis enhancements and XPath
I had dropped out of the blogging habit after Feb this year, mainly due to work pressures at office. I think it is time for me to dust out my blogging brush and start painting my blog a little bit, at least to touch up some of the loosing sheen here and there :-)
This is about a nice little utility called Gnosis utils which provides powerful XML parsing in very little code. The library is written by Dr. David Mertz who is a known authority on Python programming matters and has written some charming articles in his "Charming Python" series for IBM developerworks.
I was looking around for an XML API which provides powerful XPath parsing capabilities out of the box some time in March for a project I am doing at Spikesource. PyXML woefully lacks in this department. ElementTree, though a very good API for generic XML processing, falls short on its XPath support (no attributes etc).
Gnosis provides decent XPath support - it supports attributes, text searches but does not support attribute values, the [@attr] syntax etc.
During my work, I enhanced Gnosis XML to support some of the XPath 1.0 specs it do not support. These include,
o Support for //elem[@attr] syntax
o Search attributes by value - i.e support //elem[@attr=value] syntax
o Support //elem[last()]
o Support XPath searches in the root node
The good thing is that I have been able to extract out the Gnosis XML processing parts to a single module and add these enhancements on top of it. I plan to enhance this with the full XPath 1.0 specifications within a month or so and release it to public domain.
I hope this will address the existing void of lightweight pure Python APIs with full XPath support. The whole thing will fit inside a single module which should make it easy to use and extend.
Looking forward to completing this work!
This is about a nice little utility called Gnosis utils which provides powerful XML parsing in very little code. The library is written by Dr. David Mertz who is a known authority on Python programming matters and has written some charming articles in his "Charming Python" series for IBM developerworks.
I was looking around for an XML API which provides powerful XPath parsing capabilities out of the box some time in March for a project I am doing at Spikesource. PyXML woefully lacks in this department. ElementTree, though a very good API for generic XML processing, falls short on its XPath support (no attributes etc).
Gnosis provides decent XPath support - it supports attributes, text searches but does not support attribute values, the [@attr] syntax etc.
During my work, I enhanced Gnosis XML to support some of the XPath 1.0 specs it do not support. These include,
o Support for //elem[@attr] syntax
o Search attributes by value - i.e support //elem[@attr=value] syntax
o Support //elem[last()]
o Support XPath searches in the root node
The good thing is that I have been able to extract out the Gnosis XML processing parts to a single module and add these enhancements on top of it. I plan to enhance this with the full XPath 1.0 specifications within a month or so and release it to public domain.
I hope this will address the existing void of lightweight pure Python APIs with full XPath support. The whole thing will fit inside a single module which should make it easy to use and extend.
Looking forward to completing this work!
Friday, February 24, 2006
BangPypers meet this Sunday
After a long time, the BangPypers are meeting this Sunday, the 26th again. The meeting is organized in the Spikesource conference room named Ruby. How is that for some co-operation between two popular dynamic languages !
Sidharth Kuruvila is giving a talk on one of his favorite topics, namely generators in Python. Sridhar Ratna of Amazon will talk on TurboGears, his favorite Python web dev framework. Also, Anush Shetty, a very active Python community member in Bangalore shall organize a BoF session on Python Education.
It is going to be an exciting event. If you are interested in Pythonic happenings in India's silicon valley, make sure you drop in to Spikesource offices in Diamond District on Airport Road on the 26th at 4.30 pm sharp.
Unfortunately, I cannot attend this meet since I am going out of Bangalore this week-end to visit my family. Deepan from Spikesource should take care of the meeting co-ordination.
Hope you all have a fun and informative session!
See the meeting description in the group's calendar for details and contact information.
Sidharth Kuruvila is giving a talk on one of his favorite topics, namely generators in Python. Sridhar Ratna of Amazon will talk on TurboGears, his favorite Python web dev framework. Also, Anush Shetty, a very active Python community member in Bangalore shall organize a BoF session on Python Education.
It is going to be an exciting event. If you are interested in Pythonic happenings in India's silicon valley, make sure you drop in to Spikesource offices in Diamond District on Airport Road on the 26th at 4.30 pm sharp.
Unfortunately, I cannot attend this meet since I am going out of Bangalore this week-end to visit my family. Deepan from Spikesource should take care of the meeting co-ordination.
Hope you all have a fun and informative session!
See the meeting description in the group's calendar for details and contact information.
Friday, December 16, 2005
New exception model in IronPython
I am a lurker in the IronPython mailing lists. Now that the project is slowly approaching 1.0 release, I have started taking an active interest in the proceedings in the list.
The recent postings on the proposed new exception handling system in IronPython makes interesting reading. Apparently the current model does not provide a nice co-existence of the exception hierarchies of CPython and the .NET CLI. This means that if you call a piece of pure CLI code from inside your Python (IronPython) code and if it raises a CLI exception, you are left stranded in no-man's land if your code only catches Python exceptions.
The proposed new model plans to separate the exception hierarchy into two hieararchies - one for .NET CLI exceptions and one for Python exceptions. The Python exception hierarchy will mirror that of the existing CPython exception class hierarchy. Under this model pure Python and CLI code catching their own exceptions will work the same. When the two worlds collide, i.e when pure Python code tries to catch a CLI exception or when pure CLI code tries to catch a Python exception, the system will convert the exception to the required type.
For example, if the Python code raises EOFError, the CLI code will see an EndOfStreamException. If the CLI code raises FileIOException, the Python code will see an IOError. The error mappings happen automatically.
Some problems still remain unsolved - like what the user should see as an exception traceback; should it be pure Python exceptions or .NET ? The current thinking is that the language should allow both forms in a user-configurable way.
One issue with this model is that in order to map CLI exceptions to Python exceptions you need to go through some kind of mapping for CLI to Python exceptions; this could slow things down a bit. Also, the model requires rethrowing of exceptions whenever the CLI and Python worlds collide.
It will be interesting to wait and watch how the new model develops.
The new exception model is described in detail in this Channel 9 wiki posting.
The recent postings on the proposed new exception handling system in IronPython makes interesting reading. Apparently the current model does not provide a nice co-existence of the exception hierarchies of CPython and the .NET CLI. This means that if you call a piece of pure CLI code from inside your Python (IronPython) code and if it raises a CLI exception, you are left stranded in no-man's land if your code only catches Python exceptions.
The proposed new model plans to separate the exception hierarchy into two hieararchies - one for .NET CLI exceptions and one for Python exceptions. The Python exception hierarchy will mirror that of the existing CPython exception class hierarchy. Under this model pure Python and CLI code catching their own exceptions will work the same. When the two worlds collide, i.e when pure Python code tries to catch a CLI exception or when pure CLI code tries to catch a Python exception, the system will convert the exception to the required type.
For example, if the Python code raises EOFError, the CLI code will see an EndOfStreamException. If the CLI code raises FileIOException, the Python code will see an IOError. The error mappings happen automatically.
Some problems still remain unsolved - like what the user should see as an exception traceback; should it be pure Python exceptions or .NET ? The current thinking is that the language should allow both forms in a user-configurable way.
One issue with this model is that in order to map CLI exceptions to Python exceptions you need to go through some kind of mapping for CLI to Python exceptions; this could slow things down a bit. Also, the model requires rethrowing of exceptions whenever the CLI and Python worlds collide.
It will be interesting to wait and watch how the new model develops.
The new exception model is described in detail in this Channel 9 wiki posting.
Friday, December 02, 2005
Back to Bangalore
I came back to Bangalore after six and a half fine days in Norway on 1st early morning, 2 am. The flight from Frankfurt to Bangalore took 8 hours.
The photo album is at Y! which is online here.
After spending six days at near and below zero temperatures in Norway, Bangalore is feeling hot.
The photo album is at Y! which is online here.
After spending six days at near and below zero temperatures in Norway, Bangalore is feeling hot.
Wednesday, November 23, 2005
Workshop in Norway
I am leaving tonight (Nov 24 2005, 2.15 am) to Grimstad, Norway for taking part in the 2nd international EIAO conference and workshop. This conference is organised by AUC as part of the EIAO project.
I will be making a small presentation on D-HarvestMan. I am returning to India on 1st Dec.
Here is a brief abstract of the proceedings in Grimstad.
2nd international EIAO conference and worskhop
I will be making a small presentation on D-HarvestMan. I am returning to India on 1st Dec.
Here is a brief abstract of the proceedings in Grimstad.
2nd international EIAO conference and worskhop
24.11.05
1. Presentation of overall timeplan
2. Presentation of preliminary Observatory anatomy
3. WAM coordination (D3.1.1, UWEM, D5.1.1.1 and D5.1.1.2 )
4. Coordination of WP3, WP4 and WP5 release 1.0 time plans
5. Overall release 1.0 planning and actions
25.11.05
1. Welcome Mikael Snaprud
2. D-Crawler - Anand B. Pillai
3. Presentation of anatomy methodology Alf Fredvik
4. SW development process Design and test. Parastoo Mohagheghi,
Per Wollebæk
26.11.05 and 27.11.05
1. Presentation of overall timeplan
2. Presentation of preliminary Observatory anatomy
3. Development of a DW anatomy
4. Description of each anatomy
5. Scheduling of each anatomy
6. SW development process (Design and test)
7. Coordination of WP6 and WP5 release 1.0 time plans
8. Overall release 1.0 planning and actions
28.11.05
Follow ups and documentation
1. Inspection and adjustment of plans for each anatomy
2. Strategy for outreach and dissemination towards
and after release 1.0
3. Summary and outlook - HarvestMan as a vehicle for
research based teaching
Thursday, November 17, 2005
Internet Summit in Tunis
The World Summit on Information Society (WSIS), dubbed as the "Internet Showdown" by CNET, is being held under the aegis of the U.N in Tunis, the capital of Tunisia at present.
Dr.Mikael Snaprud of the EIAO project presented a paper in the conference(Past, Presence and Future of Research in the Information Society) associated to this summit on Nov 15. The paper talks about the role open source plays in ICT education and research, with the EIAO project as the background. The presentation associated with this paper is available at the EIAO Publications website.
The paper is authored by Mikael Snaprud, Agata Sawicka, Anand Pillai(myself), Nina Olsen, Morten.G.Olsen, Vidar Laupsa and Terje Gjøsæter.
Dr.Mikael Snaprud of the EIAO project presented a paper in the conference(Past, Presence and Future of Research in the Information Society) associated to this summit on Nov 15. The paper talks about the role open source plays in ICT education and research, with the EIAO project as the background. The presentation associated with this paper is available at the EIAO Publications website.
The paper is authored by Mikael Snaprud, Agata Sawicka, Anand Pillai(myself), Nina Olsen, Morten.G.Olsen, Vidar Laupsa and Terje Gjøsæter.
Monday, November 14, 2005
D-HarvestMan prototype is born
It is 2.30 am in the morning right now. I am in a good mood. The reason is that I just finished coding and testing the basic D-HarvestMan prototype for a single master, single slave configuration. And it works! The master was able to successfully bootstrap the slave crawler with a new domain and let it start downloading files from it. Hip hip hooray!
D-HarvestMan is a project to write a distributed crawler on top of the existing HarvestMan. Distributed programming is always exciting, and a distributed crawler is even more so :-)
D-HarvestMan is a project to write a distributed crawler on top of the existing HarvestMan. Distributed programming is always exciting, and a distributed crawler is even more so :-)
Friday, November 11, 2005
The fox is one year old
Firefox turned one year old on Nov 9, two days back. The wily fox has set the web on fire, ever since it debuted on Nov 9 2004. With the 1.5 final release on the way, it is well on course to capture further market share from the beleaguered I.E .
Three cheers to the Firefox team and wish all the best to the fox for its second year!
Three cheers to the Firefox team and wish all the best to the fox for its second year!
Monday, November 07, 2005
OOo plugin for Firefox
In the context of my previous post, it is interesting to see that people have already talked about how an OOo plugin for Firefox can be a killer collaboration application on par with MS Sharepoint.
Apparently a Mozilla plugin is already available in OOo 2.0. The problem with it is that you enable the plugin from inside OOo, which requires that OOo is already installed in your machine.
From the Mozilla OOo plugin specification,
This reduces the usability of the plugin, I think. Developing a light-weight OOo plugin for Mozilla/Firefox which can be installed on the fly could be pretty useful.
Apparently a Mozilla plugin is already available in OOo 2.0. The problem with it is that you enable the plugin from inside OOo, which requires that OOo is already installed in your machine.
From the Mozilla OOo plugin specification,
"Plug-in usability
The plug-in works only if a working OpenOffice.org installation is found on the system""
This reduces the usability of the plugin, I think. Developing a light-weight OOo plugin for Mozilla/Firefox which can be installed on the fly could be pretty useful.
Saturday, November 05, 2005
Some random thoughts on a web-based office suite
Ever since Sun and Google announced their partnership early in October, speculation has been rife on the possibility of a web-based office suite, aptly titled "GoogleOffice". However, there has not been any strong indication of such an effort underway. In fact a number of industry watchers were disappointed when Google and Sun announced that the initial collaboration would be on bundling the Google toolbar with Java runtime downloads.
Now that Microsoft is alligning itself as a services provider and trying to offer integrated solutions by bundling its diverse product portfolio (Office, MSN, messenger etc) along with its new service initiatives (Windows Live, Office Live), it is probably time that Google looked into countering these overtures with appropriate answers - I think there is nothing more fitting here than an office solution which integrates Gmail, Google Talk, Google Desktop and the Firefox web browser.
Perhaps such an effort is already underway in Google Labs. However, here is my vision for such a solution, a kind of blue print for a future GoogleOffice on the web.
Google office will be a collaborative suite integrating Gmail, Google Desktop, Google Talk and Firefox. Ideally it should have the following components:
1. An open-office plugin for Firefox
2. An extension to Google desktop that allows Gmail attachments to be searched and accessed.
3. An extension to Google talk that allows members to access attachments in their Gmail account and also share files and folders through Google Desktop.
Let me explain:
1. The Firefox plugin will allow one to view Openoffice documents inside the Firefox web-browser. Initially this need to support only OOo native file formats and the OpenDocument format, but MS office support would be preferable.
2. Right now attachments cannot be searched in Gmail. Google needs to add this capability to Gmail. Attachments can be searched by name, but they also need to be searchable by content.
3. Google desktop integrates with Gmail, but again attachments are not searchable or accessible. This capability need to be added to Google Desktop so that one can search and access documents stored as attachments in his Gmail id through Google Desktop.
4. The same capability should be added to Google Talk so that one can search for attached documents from Google Talk.
5. Integration between Google Talk and Google Desktop so that chatters can share documents with each other and search each others desktop, given sufficient access control privileges.
All these pieces will allow for a basic GoogleOffice over the web. Let us look at some common scenarios:
1. Someone is browsing in an Internet cafe - He wants to view an OOo document sent to his Gmail id as email attachment. Typically Internet cafes do not have OOo so he is at a loss (This happens to me quite often). This is where an OOo plugin for Firefox can help. Gmail need not know anything about this plugin. It can be a regular firefox plugin. In this case, Firefox will detect that the plugin is not installed, will download and install it automatically. Voila, you can view your OOo document inside firefox.
I dont think it will be too difficult to develop such a plugin, considering that the source code for OOo is open and the OOo file formats are well documented. This will also help in large scale acceptance of OOo file formats in a way similar to PDF. Editing capabilities will also be nice but this could be tough to implement in a browser plugin.
I would like to see plugins for all OOo file formats but specifically OOo writer (.sxw), OOo impress (.sxi) and the OpenDocument formats.
2. Adding capability to search and find attachments in Gmail will allow people to use Gmail as a sort of virtual storage for their documents (read office documents). I tend to do this even now, but because the attachments are not searchable, the experience is crippled. Initially these can be added for well known and open formats such as PDF and the OOo file formats.
3. Once Gmail is enhanced with advanced capabilities to search attachments, this capability should be integrated with Google desktop which can then index the attachments and make them searchable from the desktop. Considering the work involved in indexing large attachments, it actually makes sense to add this capabillity only to Google desktop rather than onto Gmail directly.
4. This opens up the possibility of integrating Firefox, Gmail and Google Desktop for searching, accessing and modifying office documents. One can use Google Desktop to search office doucments stored in his Gmail account, then open and edit them on his desktop inside Firefox using the OOo plugin. If OOo is installed on the machine already, it can be used instead.
5. The last missing piece is Google Talk. Google Talk should integrate with Google desktop allowing querying of documents (Gmail, desktop, photos etc) through Google Talk. Google should add capabilities to Google Talk which will make this possible not only on one's own desktop but across desktops.
That is, you can give privilges in Google Talk to selected Google Talk users (your friends, co-workers, family), to search and access documents from inside your desktop and also your Gmail id. These can be controlled by access at various levels - read-only, read-write etc.
If one finds documents of interest they can be shared and edited online. Google should provide something like a shared space for two interested parties to share documents as a part of Google talk. This space can either be part of Gmail or separate from it. However, it should allow for saving documents by multiple Google talk users, acting as a kind of collaborative space.
Thus Google Talk and Google desktop along with Gmail and Firefox can be used to build a virtual collaborative office suite on the web which if done well, can probably pose some compeition to Microsoft office solutions and the recent office collaboration initiatives. It will also allow to take the OOo efforts to the web and provide it as a part of a services offering instead of the current stand-alone product.
Perhaps this vision is a bit grand, but I don't think it is a difficult one for Google and the open source (read Openoffice) community. All the pieces are already there, they just require some additional capabilities and some plumbing to work as a unified, web-based solution.
I have not thought to deep about the technical aspects of such a solution but a very interesting thought will be the role Java and Google toolbar can play in this integrated approach. Perhaps Java can be used to develop the Firefox OOo plugin also.
I hope we can expect to see a web-based office solution from Google, Sun and the OOo community within the next 12 months.
Now that Microsoft is alligning itself as a services provider and trying to offer integrated solutions by bundling its diverse product portfolio (Office, MSN, messenger etc) along with its new service initiatives (Windows Live, Office Live), it is probably time that Google looked into countering these overtures with appropriate answers - I think there is nothing more fitting here than an office solution which integrates Gmail, Google Talk, Google Desktop and the Firefox web browser.
Perhaps such an effort is already underway in Google Labs. However, here is my vision for such a solution, a kind of blue print for a future GoogleOffice on the web.
Google office will be a collaborative suite integrating Gmail, Google Desktop, Google Talk and Firefox. Ideally it should have the following components:
1. An open-office plugin for Firefox
2. An extension to Google desktop that allows Gmail attachments to be searched and accessed.
3. An extension to Google talk that allows members to access attachments in their Gmail account and also share files and folders through Google Desktop.
Let me explain:
1. The Firefox plugin will allow one to view Openoffice documents inside the Firefox web-browser. Initially this need to support only OOo native file formats and the OpenDocument format, but MS office support would be preferable.
2. Right now attachments cannot be searched in Gmail. Google needs to add this capability to Gmail. Attachments can be searched by name, but they also need to be searchable by content.
3. Google desktop integrates with Gmail, but again attachments are not searchable or accessible. This capability need to be added to Google Desktop so that one can search and access documents stored as attachments in his Gmail id through Google Desktop.
4. The same capability should be added to Google Talk so that one can search for attached documents from Google Talk.
5. Integration between Google Talk and Google Desktop so that chatters can share documents with each other and search each others desktop, given sufficient access control privileges.
All these pieces will allow for a basic GoogleOffice over the web. Let us look at some common scenarios:
1. Someone is browsing in an Internet cafe - He wants to view an OOo document sent to his Gmail id as email attachment. Typically Internet cafes do not have OOo so he is at a loss (This happens to me quite often). This is where an OOo plugin for Firefox can help. Gmail need not know anything about this plugin. It can be a regular firefox plugin. In this case, Firefox will detect that the plugin is not installed, will download and install it automatically. Voila, you can view your OOo document inside firefox.
I dont think it will be too difficult to develop such a plugin, considering that the source code for OOo is open and the OOo file formats are well documented. This will also help in large scale acceptance of OOo file formats in a way similar to PDF. Editing capabilities will also be nice but this could be tough to implement in a browser plugin.
I would like to see plugins for all OOo file formats but specifically OOo writer (.sxw), OOo impress (.sxi) and the OpenDocument formats.
2. Adding capability to search and find attachments in Gmail will allow people to use Gmail as a sort of virtual storage for their documents (read office documents). I tend to do this even now, but because the attachments are not searchable, the experience is crippled. Initially these can be added for well known and open formats such as PDF and the OOo file formats.
3. Once Gmail is enhanced with advanced capabilities to search attachments, this capability should be integrated with Google desktop which can then index the attachments and make them searchable from the desktop. Considering the work involved in indexing large attachments, it actually makes sense to add this capabillity only to Google desktop rather than onto Gmail directly.
4. This opens up the possibility of integrating Firefox, Gmail and Google Desktop for searching, accessing and modifying office documents. One can use Google Desktop to search office doucments stored in his Gmail account, then open and edit them on his desktop inside Firefox using the OOo plugin. If OOo is installed on the machine already, it can be used instead.
5. The last missing piece is Google Talk. Google Talk should integrate with Google desktop allowing querying of documents (Gmail, desktop, photos etc) through Google Talk. Google should add capabilities to Google Talk which will make this possible not only on one's own desktop but across desktops.
That is, you can give privilges in Google Talk to selected Google Talk users (your friends, co-workers, family), to search and access documents from inside your desktop and also your Gmail id. These can be controlled by access at various levels - read-only, read-write etc.
If one finds documents of interest they can be shared and edited online. Google should provide something like a shared space for two interested parties to share documents as a part of Google talk. This space can either be part of Gmail or separate from it. However, it should allow for saving documents by multiple Google talk users, acting as a kind of collaborative space.
Thus Google Talk and Google desktop along with Gmail and Firefox can be used to build a virtual collaborative office suite on the web which if done well, can probably pose some compeition to Microsoft office solutions and the recent office collaboration initiatives. It will also allow to take the OOo efforts to the web and provide it as a part of a services offering instead of the current stand-alone product.
Perhaps this vision is a bit grand, but I don't think it is a difficult one for Google and the open source (read Openoffice) community. All the pieces are already there, they just require some additional capabilities and some plumbing to work as a unified, web-based solution.
I have not thought to deep about the technical aspects of such a solution but a very interesting thought will be the role Java and Google toolbar can play in this integrated approach. Perhaps Java can be used to develop the Firefox OOo plugin also.
I hope we can expect to see a web-based office solution from Google, Sun and the OOo community within the next 12 months.
Wednesday, November 02, 2005
FOSS.in
FOSS.in has published the second list of speakers of the event. Some notables who are speaking include the legendary Alan Cox, Danese Cooper, David Fetter, Jeremy Zawodny, Jonathan Corbet, Andrew Cowie, Harald Welte and Brian Behlendorf.
Murugan Pal is giving a talk on "Open source Alternatives".
Also the talk by Zaheda Bhorat of Google on Google and open source seems interesting.
FOSS.in is scheduled from 29th Nov to 2nd Dec 2005 at Bangalore.
Murugan Pal is giving a talk on "Open source Alternatives".
Also the talk by Zaheda Bhorat of Google on Google and open source seems interesting.
FOSS.in is scheduled from 29th Nov to 2nd Dec 2005 at Bangalore.
On Open Voting
An article published in a local daily of Granite Bay, CA, talks about accountability and the Open Voting Consortium.
Read the article.
What is my interest in this ? Well, I happen to be part of the team that originally developed the OVC prototype, which was demonstrated in April 1 2004. The OVC project was my first experience in working for an international open source project. The architects of the system decided to use Python for the project, which was how I got interested in it.
I am also a founding member of the OVC.
Read the article.
What is my interest in this ? Well, I happen to be part of the team that originally developed the OVC prototype, which was demonstrated in April 1 2004. The OVC project was my first experience in working for an international open source project. The architects of the system decided to use Python for the project, which was how I got interested in it.
I am also a founding member of the OVC.
Saturday, October 22, 2005
International Conference on Digital Inclusion and Open Source
The annual conference on "Digal Inclusion and Open Source" took place in Oslo, Norway from Oct 20-21, 2005.
A paper co-authored by me, Parastoo Mohaghegi, Mikael Snaprud & Nils Ulltveit-Moe was presented in the conference. The paper highlights the activities of the EIAO project in bringing together users and external contributors from different parts of the world in an EU sponsored project.
Read the abstract of the paper.
A paper co-authored by me, Parastoo Mohaghegi, Mikael Snaprud & Nils Ulltveit-Moe was presented in the conference. The paper highlights the activities of the EIAO project in bringing together users and external contributors from different parts of the world in an EU sponsored project.
Read the abstract of the paper.
Tuesday, October 04, 2005
HarvestMan web-site redesign
I have finally re-designed the HarvestMan web-site! It has been something I have been planning since the beginning of this year! When I finally did it, it took me only two days. Surprising, how much savings one can get in terms of time, if one really focuses on the task at hand.
Wednesday, September 28, 2005
Microsoft and JBoss shake hands
Probably bad news for LAMP and other open source stacks. Complete news is here.
Tuesday, September 13, 2005
Saturday, September 10, 2005
Cathedral tries to recruit Bazaar!
No kidding :-). Microsoft apparently tried to recruit Eric Raymond. If you don't know who Eric Raymond is, spend your afternoon reading up the excellent Cathedral and Bazaar essays, some of the best essays written on the open source model. He also happens to be one of the co-founders of OSI. Talk about Bush trying to recruit Osama Bin Laden for homeland security!
Read more about Microsoft's Mea Culpa in this article posted on ESR's blog.
Read more about Microsoft's Mea Culpa in this article posted on ESR's blog.
Friday, September 09, 2005
Standalone Executables on Windows using py2exe
The latest release of py2exe, namely py2exe 0.6.1 allows to create single executables on Windows. This is an improvement over the earlier versions which used to create a host of files around the main executable. I think the new feature is a welcome one, especially for a project like HarvestMan which has a number of dependencies. If you try to create an executable for HarvestMan with existing versions
of py2exe, you get quite a lot of .pyd files which are well, a bit confusing.
I am looking forward to create standalone executables for HarvestMan using the latest py2exe and provide downloads for them. I think this should boost the popularity of HarvestMan, since many Windows users I know could not be bothered with going through all the steps to install a pure Python package such as HarvestMan. Downloading and installing a single file executable is much easier.
Expect win32 downloads of HarvestMan soon. Three cheers to py2exe and Thomas Heller.
of py2exe, you get quite a lot of .pyd files which are well, a bit confusing.
I am looking forward to create standalone executables for HarvestMan using the latest py2exe and provide downloads for them. I think this should boost the popularity of HarvestMan, since many Windows users I know could not be bothered with going through all the steps to install a pure Python package such as HarvestMan. Downloading and installing a single file executable is much easier.
Expect win32 downloads of HarvestMan soon. Three cheers to py2exe and Thomas Heller.
Subscribe to:
Posts (Atom)
