# Edmund Miller Bioinformatics, biology, genomics, and software engineering notes from Edmund Miller. ## Posts --- # Richard Hamming: You and Your Research Richard Hamming on choosing important problems, doing first-class work, and making the most of the one life we have. import { YouTube } from '@astro-community/astro-embed-youtube'; Mathys shared this talk with me, and it has been one of the most influential things I have read. I keep coming back to the question it taught me to ask: > What is the most important thing I could be working on right now? Richard Hamming delivered “You and Your Research” at Bellcore on March 7, 1986. [Paul Graham hosts the copy I first read](https://www.paulgraham.com/hamming.html). I am preserving it here because I want to keep returning to it. ## Watch the talk This is a later version of “You and Your Research,” recorded on June 6, 1995. --- _Talk at Bellcore, 7 March 1986_ The title of my talk is "You and Your Research." It is not about managing research, it is about how you individually do your research. I could give a talk on the other subject — but it's not, it's about you. I'm not talking about ordinary run-of-the-mill research; I'm talking about great research. And for the sake of describing great research I'll occasionally say Nobel-Prize type of work. It doesn't have to gain the Nobel Prize, but I mean those kinds of things which we perceive are significant things. Relativity, if you want, Shannon's information theory, any number of outstanding theories — that's the kind of thing I'm talking about. Now, how did I come to do this study? At Los Alamos I was brought in to run the computing machines which other people had got going, so those scientists and physicists could get back to business. I saw I was a stooge. I saw that although physically I was the same, they were different. And to put the thing bluntly, I was envious. I wanted to know why they were so different from me. I saw Feynman up close. I saw Fermi and Teller. I saw Oppenheimer. I saw Hans Bethe: he was my boss. I saw quite a few very capable people. I became very interested in the difference between those who do and those who might have done. When I came to Bell Labs, I came into a very productive department. Bode was the department head at the time; Shannon was there, and there were other people. I continued examining the questions, "Why?" and "What is the difference?" I continued subsequently by reading biographies, autobiographies, asking people questions such as: "How did you come to do this?" I tried to find out what are the differences. And that's what this talk is about. Now, why is this talk important? I think it is important because, as far as I know, each of you has one life to live. Even if you believe in reincarnation it doesn't do you any good from one life to the next! Why shouldn't you do significant things in this one life, however you define significant? I'm not going to define it — you know what I mean. I will talk mainly about science because that is what I have studied. But so far as I know, and I've been told by others, much of what I say applies to many fields. Outstanding work is characterized very much the same way in most fields, but I will confine myself to science. In order to get at you individually, I must talk in the first person. I have to get you to drop modesty and say to yourself, "Yes, I would like to do first-class work." Our society frowns on people who set out to do really good work. You're not supposed to; luck is supposed to descend on you and you do great things by chance. Well, that's a kind of dumb thing to say. I say, why shouldn't you set out to do something significant. You don't have to tell other people, but shouldn't you say to yourself, "Yes, I would like to do something significant." In order to get to the second stage, I have to drop modesty and talk in the first person about what I've seen, what I've done, and what I've heard. I'm going to talk about people, some of whom you know, and I trust that when we leave, you won't quote me as saying some of the things I said. Let me start not logically, but psychologically. I find that the major objection is that people think great science is done by luck. It's all a matter of luck. Well, consider Einstein. Note how many different things he did that were good. Was it all luck? Wasn't it a little too repetitive? Consider Shannon. He didn't do just information theory. Several years before, he did some other good things and some which are still locked up in the security of cryptography. He did many good things. You see again and again that it is more than one thing from a good person. Once in a while a person does only one thing in his whole life, and we'll talk about that later, but a lot of times there is repetition. I claim that luck will not cover everything. And I will cite Pasteur who said, "Luck favors the prepared mind." And I think that says it the way I believe it. There is indeed an element of luck, and no, there isn't. The prepared mind sooner or later finds something important and does it. So yes, it is luck. The particular thing you do is luck, but that you do something is not. For example, when I came to Bell Labs, I shared an office for a while with Shannon. At the same time he was doing information theory, I was doing coding theory. It is suspicious that the two of us did it at the same place and at the same time — it was in the atmosphere. And you can say, "Yes, it was luck." On the other hand you can say, "But why of all the people in Bell Labs then were those the two who did it?" Yes, it is partly luck, and partly it is the prepared mind; but "partly" is the other thing I'm going to talk about. So, although I'll come back several more times to luck, I want to dispose of this matter of luck as being the sole criterion whether you do great work or not. I claim you have some, but not total, control over it. And I will quote, finally, Newton on the matter. Newton said, "If others would think as hard as I did, then they would get similar results." One of the characteristics you see, and many people have it including great scientists, is that usually when they were young they had independent thoughts and had the courage to pursue them. For example, Einstein, somewhere around 12 or 14, asked himself the question, "What would a light wave look like if I went with the velocity of light to look at it?" Now he knew that electromagnetic theory says you cannot have a stationary local maximum. But if he moved along with the velocity of light, he would see a local maximum. He could see a contradiction at the age of 12, 14, or somewhere around there, that everything was not right and that the velocity of light had something peculiar. Is it luck that he finally created special relativity? Early on, he had laid down some of the pieces by thinking of the fragments. Now that's the necessary but not sufficient condition. All of these items I will talk about are both luck and not luck. How about having lots of brains? It sounds good. Most of you in this room probably have more than enough brains to do first-class work. But great work is something else than mere brains. Brains are measured in various ways. In mathematics, theoretical physics, astrophysics, typically brains correlates to a great extent with the ability to manipulate symbols. And so the typical IQ test is apt to score them fairly high. On the other hand, in other fields it is something different. For example, Bill Pfann, the fellow who did zone melting, came into my office one day. He had this idea dimly in his mind about what he wanted and he had some equations. It was pretty clear to me that this man didn't know much mathematics and he wasn't really articulate. His problem seemed interesting so I took it home and did a little work. I finally showed him how to run computers so he could compute his own answers. I gave him the power to compute. He went ahead, with negligible recognition from his own department, but ultimately he has collected all the prizes in the field. Once he got well started, his shyness, his awkwardness, his inarticulateness, fell away and he became much more productive in many other ways. Certainly he became much more articulate. And I can cite another person in the same way. I trust he isn't in the audience, i.e. a fellow named Clogston. I met him when I was working on a problem with John Pierce's group and I didn't think he had much. I asked my friends who had been with him at school, "Was he like that in graduate school?" "Yes," they replied. Well I would have fired the fellow, but J. R. Pierce was smart and kept him on. Clogston finally did the Clogston cable. After that there was a steady stream of good ideas. One success brought him confidence and courage. One of the characteristics of successful scientists is having courage. Once you get your courage up and believe that you can do important problems, then you can. If you think you can't, almost surely you are not going to. Courage is one of the things that Shannon had supremely. You have only to think of his major theorem. He wants to create a method of coding, but he doesn't know what to do so he makes a random code. Then he is stuck. And then he asks the impossible question, "What would the average random code do?" He then proves that the average code is arbitrarily good, and that therefore there must be at least one good code. Who but a man of infinite courage could have dared to think those thoughts? That is the characteristic of great scientists; they have courage. They will go forward under incredible circumstances; they think and continue to think. Age is another factor which the physicists particularly worry about. They always are saying that you have got to do it when you are young or you will never do it. Einstein did things very early, and all the quantum mechanic fellows were disgustingly young when they did their best work. Most mathematicians, theoretical physicists, and astrophysicists do what we consider their best work when they are young. It is not that they don't do good work in their old age but what we value most is often what they did early. On the other hand, in music, politics and literature, often what we consider their best work was done late. I don't know how whatever field you are in fits this scale, but age has some effect. But let me say why age seems to have the effect it does. In the first place if you do some good work you will find yourself on all kinds of committees and unable to do any more work. You may find yourself as I saw Brattain when he got a Nobel Prize. The day the prize was announced we all assembled in Arnold Auditorium; all three winners got up and made speeches. The third one, Brattain, practically with tears in his eyes, said, "I know about this Nobel-Prize effect and I am not going to let it affect me; I am going to remain good old Walter Brattain." Well I said to myself, "That is nice." But in a few weeks I saw it was affecting him. Now he could only work on great problems. When you are famous it is hard to work on small problems. This is what did Shannon in. After information theory, what do you do for an encore? The great scientists often make this error. They fail to continue to plant the little acorns from which the mighty oak trees grow. They try to get the big thing right off. And that isn't the way things go. So that is another reason why you find that when you get early recognition it seems to sterilize you. In fact I will give you my favorite quotation of many years. The Institute for Advanced Study in Princeton, in my opinion, has ruined more good scientists than any institution has created, judged by what they did before they came and judged by what they did after. Not that they weren't good afterwards, but they were superb before they got there and were only good afterwards. This brings up the subject, out of order perhaps, of working conditions. What most people think are the best working conditions, are not. Very clearly they are not because people are often most productive when working conditions are bad. One of the better times of the Cambridge Physical Laboratories was when they had practically shacks — they did some of the best physics ever. I give you a story from my own private life. Early on it became evident to me that Bell Laboratories was not going to give me the conventional acre of programming people to program computing machines in absolute binary. It was clear they weren't going to. But that was the way everybody did it. I could go to the West Coast and get a job with the airplane companies without any trouble, but the exciting people were at Bell Labs and the fellows out there in the airplane companies were not. I thought for a long while about, "Did I want to go or not?" and I wondered how I could get the best of two possible worlds. I finally said to myself, "Hamming, you think the machines can do practically everything. Why can't you make them write programs?" What appeared at first to me as a defect forced me into automatic programming very early. What appears to be a fault, often, by a change of viewpoint, turns out to be one of the greatest assets you can have. But you are not likely to think that when you first look the thing and say, "Gee, I'm never going to get enough programmers, so how can I ever do any great programming?" And there are many other stories of the same kind; Grace Hopper has similar ones. I think that if you look carefully you will see that often the great scientists, by turning the problem around a bit, changed a defect to an asset. For example, many scientists when they found they couldn't do a problem finally began to study why not. They then turned it around the other way and said, "But of course, this is what it is" and got an important result. So ideal working conditions are very strange. The ones you want aren't always the best ones for you. Now for the matter of drive. You observe that most great scientists have tremendous drive. I worked for ten years with John Tukey at Bell Labs. He had tremendous drive. One day about three or four years after I joined, I discovered that John Tukey was slightly younger than I was. John was a genius and I clearly was not. Well I went storming into Bode's office and said, "How can anybody my age know as much as John Tukey does?" He leaned back in his chair, put his hands behind his head, grinned slightly, and said, "You would be surprised Hamming, how much you would know if you worked as hard as he did that many years." I simply slunk out of the office! What Bode was saying was this: Knowledge and productivity are like compound interest. Given two people of approximately the same ability and one person who works ten percent more than the other, the latter will more than twice outproduce the former. The more you know, the more you learn; the more you learn, the more you can do; the more you can do, the more the opportunity — it is very much like compound interest. I don't want to give you a rate, but it is a very high rate. Given two people with exactly the same ability, the one person who manages day in and day out to get in one more hour of thinking will be tremendously more productive over a lifetime. I took Bode's remark to heart; I spent a good deal more of my time for some years trying to work a bit harder and I found, in fact, I could get more work done. I don't like to say it in front of my wife, but I did sort of neglect her sometimes; I needed to study. You have to neglect things if you intend to get what you want done. There's no question about this. On this matter of drive Edison says, "Genius is 99% perspiration and 1% inspiration." He may have been exaggerating, but the idea is that solid work, steadily applied, gets you surprisingly far. The steady application of effort with a little bit more work, intelligently applied is what does it. That's the trouble; drive, misapplied, doesn't get you anywhere. I've often wondered why so many of my good friends at Bell Labs who worked as hard or harder than I did, didn't have so much to show for it. The misapplication of effort is a very serious matter. Just hard work is not enough - it must be applied sensibly. There's another trait on the side which I want to talk about; that trait is ambiguity. It took me a while to discover its importance. Most people like to believe something is or is not true. Great scientists tolerate ambiguity very well. They believe the theory enough to go ahead; they doubt it enough to notice the errors and faults so they can step forward and create the new replacement theory. If you believe too much you'll never notice the flaws; if you doubt too much you won't get started. It requires a lovely balance. But most great scientists are well aware of why their theories are true and they are also well aware of some slight misfits which don't quite fit and they don't forget it. Darwin writes in his autobiography that he found it necessary to write down every piece of evidence which appeared to contradict his beliefs because otherwise they would disappear from his mind. When you find apparent flaws you've got to be sensitive and keep track of those things, and keep an eye out for how they can be explained or how the theory can be changed to fit them. Those are often the great contributions. Great contributions are rarely done by adding another decimal place. It comes down to an emotional commitment. Most great scientists are completely committed to their problem. Those who don't become committed seldom produce outstanding, first-class work. Now again, emotional commitment is not enough. It is a necessary condition apparently. And I think I can tell you the reason why. Everybody who has studied creativity is driven finally to saying, "creativity comes out of your subconscious." Somehow, suddenly, there it is. It just appears. Well, we know very little about the subconscious; but one thing you are pretty well aware of is that your dreams also come out of your subconscious. And you're aware your dreams are, to a fair extent, a reworking of the experiences of the day. If you are deeply immersed and committed to a topic, day after day after day, your subconscious has nothing to do but work on your problem. And so you wake up one morning, or on some afternoon, and there's the answer. For those who don't get committed to their current problem, the subconscious goofs off on other things and doesn't produce the big result. So the way to manage yourself is that when you have a real important problem you don't let anything else get the center of your attention — you keep your thoughts on the problem. Keep your subconscious starved so it has to work on your problem, so you can sleep peacefully and get the answer in the morning, free. Now Alan Chynoweth mentioned that I used to eat at the physics table. I had been eating with the mathematicians and I found out that I already knew a fair amount of mathematics; in fact, I wasn't learning much. The physics table was, as he said, an exciting place, but I think he exaggerated on how much I contributed. It was very interesting to listen to Shockley, Brattain, Bardeen, J. B. Johnson, Ken McKay and other people, and I was learning a lot. But unfortunately a Nobel Prize came, and a promotion came, and what was left was the dregs. Nobody wanted what was left. Well, there was no use eating with them! Over on the other side of the dining hall was a chemistry table. I had worked with one of the fellows, Dave McCall; furthermore he was courting our secretary at the time. I went over and said, "Do you mind if I join you?" They can't say no, so I started eating with them for a while. And I started asking, "What are the important problems of your field?" And after a week or so, "What important problems are you working on?" And after some more time I came in one day and said, "If what you are doing is not important, and if you don't think it is going to lead to something important, why are you at Bell Labs working on it?" I wasn't welcomed after that; I had to find somebody else to eat with! That was in the spring. In the fall, Dave McCall stopped me in the hall and said, "Hamming, that remark of yours got underneath my skin. I thought about it all summer, i.e. what were the important problems in my field. I haven't changed my research," he says, "but I think it was well worthwhile." And I said, "Thank you Dave," and went on. I noticed a couple of months later he was made the head of the department. I noticed the other day he was a Member of the National Academy of Engineering. I noticed he has succeeded. I have never heard the names of any of the other fellows at that table mentioned in science and scientific circles. They were unable to ask themselves, "What are the important problems in my field?" If you do not work on an important problem, it's unlikely you'll do important work. It's perfectly obvious. Great scientists have thought through, in a careful way, a number of important problems in their field, and they keep an eye on wondering how to attack them. Let me warn you, "important problem" must be phrased carefully. The three outstanding problems in physics, in a certain sense, were never worked on while I was at Bell Labs. By important I mean guaranteed a Nobel Prize and any sum of money you want to mention. We didn't work on (1) time travel, (2) teleportation, and (3) antigravity. They are not important problems because we do not have an attack. It's not the consequence that makes a problem important, it is that you have a reasonable attack. That is what makes a problem important. When I say that most scientists don't work on important problems, I mean it in that sense. The average scientist, so far as I can make out, spends almost all his time working on problems which he believes will not be important and he also doesn't believe that they will lead to important problems. I spoke earlier about planting acorns so that oaks will grow. You can't always know exactly where to be, but you can keep active in places where something might happen. And even if you believe that great science is a matter of luck, you can stand on a mountain top where lightning strikes; you don't have to hide in the valley where you're safe. But the average scientist does routine safe work almost all the time and so he (or she) doesn't produce much. It's that simple. If you want to do great work, you clearly must work on important problems, and you should have an idea. Along those lines at some urging from John Tukey and others, I finally adopted what I called "Great Thoughts Time." When I went to lunch Friday noon, I would only discuss great thoughts after that. By great thoughts I mean ones like: "What will be the role of computers in all of AT&T?", "How will computers change science?" For example, I came up with the observation at that time that nine out of ten experiments were done in the lab and one in ten on the computer. I made a remark to the vice presidents one time, that it would be reversed, i.e. nine out of ten experiments would be done on the computer and one in ten in the lab. They knew I was a crazy mathematician and had no sense of reality. I knew they were wrong and they've been proved wrong while I have been proved right. They built laboratories when they didn't need them. I saw that computers were transforming science because I spent a lot of time asking "What will be the impact of computers on science and how can I change it?" I asked myself, "How is it going to change Bell Labs?" I remarked one time, in the same address, that more than one-half of the people at Bell Labs will be interacting closely with computing machines before I leave. Well, you all have terminals now. I thought hard about where was my field going, where were the opportunities, and what were the important things to do. Let me go there so there is a chance I can do important things. Most great scientists know many important problems. They have something between 10 and 20 important problems for which they are looking for an attack. And when they see a new idea come up, one hears them say "Well that bears on this problem." They drop all the other things and get after it. Now I can tell you a horror story that was told to me but I can't vouch for the truth of it. I was sitting in an airport talking to a friend of mine from Los Alamos about how it was lucky that the fission experiment occurred over in Europe when it did because that got us working on the atomic bomb here in the US. He said "No; at Berkeley we had gathered a bunch of data; we didn't get around to reducing it because we were building some more equipment, but if we had reduced that data we would have found fission." They had it in their hands and they didn't pursue it. They came in second! The great scientists, when an opportunity opens up, get after it and they pursue it. They drop all other things. They get rid of other things and they get after an idea because they had already thought the thing through. Their minds are prepared; they see the opportunity and they go after it. Now of course lots of times it doesn't work out, but you don't have to hit many of them to do some great science. It's kind of easy. One of the chief tricks is to live a long time! Another trait, it took me a while to notice. I noticed the following facts about people who work with the door open or the door closed. I notice that if you have the door to your office closed, you get more work done today and tomorrow, and you are more productive than most. But 10 years later somehow you don't know quite know what problems are worth working on; all the hard work you do is sort of tangential in importance. He who works with the door open gets all kinds of interruptions, but he also occasionally gets clues as to what the world is and what might be important. Now I cannot prove the cause and effect sequence because you might say, "The closed door is symbolic of a closed mind." I don't know. But I can say there is a pretty good correlation between those who work with the doors open and those who ultimately do important things, although people who work with doors closed often work harder. Somehow they seem to work on slightly the wrong thing — not much, but enough that they miss fame. I want to talk on another topic. It is based on the song which I think many of you know, "It ain't what you do, it's the way that you do it." I'll start with an example of my own. I was conned into doing on a digital computer, in the absolute binary days, a problem which the best analog computers couldn't do. And I was getting an answer. When I thought carefully and said to myself, "You know, Hamming, you're going to have to file a report on this military job; after you spend a lot of money you're going to have to account for it and every analog installation is going to want the report to see if they can't find flaws in it." I was doing the required integration by a rather crummy method, to say the least, but I was getting the answer. And I realized that in truth the problem was not just to get the answer; it was to demonstrate for the first time, and beyond question, that I could beat the analog computer on its own ground with a digital machine. I reworked the method of solution, created a theory which was nice and elegant, and changed the way we computed the answer; the results were no different. The published report had an elegant method which was later known for years as "Hamming's Method of Integrating Differential Equations." It is somewhat obsolete now, but for a while it was a very good method. By changing the problem slightly, I did important work rather than trivial work. In the same way, when using the machine up in the attic in the early days, I was solving one problem after another after another; a fair number were successful and there were a few failures. I went home one Friday after finishing a problem, and curiously enough I wasn't happy; I was depressed. I could see life being a long sequence of one problem after another after another. After quite a while of thinking I decided, "No, I should be in the mass production of a variable product. I should be concerned with all of next year's problems, not just the one in front of my face." By changing the question I still got the same kind of results or better, but I changed things and did important work. I attacked the major problem — How do I conquer machines and do all of next year's problems when I don't know what they are going to be? How do I prepare for it? How do I do this one so I'll be on top of it? How do I obey Newton's rule? He said, "If I have seen further than others, it is because I've stood on the shoulders of giants." These days we stand on each other's feet! You should do your job in such a fashion that others can build on top of it, so they will indeed say, "Yes, I've stood on so and so's shoulders and I saw further." The essence of science is cumulative. By changing a problem slightly you can often do great work rather than merely good work. Instead of attacking isolated problems, I made the resolution that I would never again solve an isolated problem except as characteristic of a class. Now if you are much of a mathematician you know that the effort to generalize often means that the solution is simple. Often by stopping and saying, "This is the problem he wants but this is characteristic of so and so. Yes, I can attack the whole class with a far superior method than the particular one because I was earlier embedded in needless detail." The business of abstraction frequently makes things simple. Furthermore, I filed away the methods and prepared for the future problems. To end this part, I'll remind you, "It is a poor workman who blames his tools — the good man gets on with the job, given what he's got, and gets the best answer he can." And I suggest that by altering the problem, by looking at the thing differently, you can make a great deal of difference in your final productivity because you can either do it in such a fashion that people can indeed build on what you've done, or you can do it in such a fashion that the next person has to essentially duplicate again what you've done. It isn't just a matter of the job, it's the way you write the report, the way you write the paper, the whole attitude. It's just as easy to do a broad, general job as one very special case. And it's much more satisfying and rewarding! I have now come down to a topic which is very distasteful; it is not sufficient to do a job, you have to sell it. "Selling" to a scientist is an awkward thing to do. It's very ugly; you shouldn't have to do it. The world is supposed to be waiting, and when you do something great, they should rush out and welcome it. But the fact is everyone is busy with their own work. You must present it so well that they will set aside what they are doing, look at what you've done, read it, and come back and say, "Yes, that was good." I suggest that when you open a journal, as you turn the pages, you ask why you read some articles and not others. You had better write your report so when it is published in the Physical Review, or wherever else you want it, as the readers are turning the pages they won't just turn your pages but they will stop and read yours. If they don't stop and read it, you won't get credit. There are three things you have to do in selling. You have to learn to write clearly and well so that people will read it, you must learn to give reasonably formal talks, and you also must learn to give informal talks. We had a lot of so-called “back room scientists.” In a conference, they would keep quiet. Three weeks later after a decision was made they filed a report saying why you should do so and so. Well, it was too late. They would not stand up right in the middle of a hot conference, in the middle of activity, and say, "We should do this for these reasons." You need to master that form of communication as well as prepared speeches. When I first started, I got practically physically ill while giving a speech, and I was very, very nervous. I realized I either had to learn to give speeches smoothly or I would essentially partially cripple my whole career. The first time IBM asked me to give a speech in New York one evening, I decided I was going to give a really good speech, a speech that was wanted, not a technical one but a broad one, and at the end if they liked it, I'd quietly say, "Any time you want one I'll come in and give you one." As a result, I got a great deal of practice giving speeches to a limited audience and I got over being afraid. Furthermore, I could also then study what methods were effective and what were ineffective. While going to meetings I had already been studying why some papers are remembered and most are not. The technical person wants to give a highly limited technical talk. Most of the time the audience wants a broad general talk and wants much more survey and background than the speaker is willing to give. As a result, many talks are ineffective. The speaker names a topic and suddenly plunges into the details he's solved. Few people in the audience may follow. You should paint a general picture to say why it's important, and then slowly give a sketch of what was done. Then a larger number of people will say, "Yes, Joe has done that," or "Mary has done that; I really see where it is; yes, Mary really gave a good talk; I understand what Mary has done." The tendency is to give a highly restricted, safe talk; this is usually ineffective. Furthermore, many talks are filled with far too much information. So I say this idea of selling is obvious. Let me summarize. You've got to work on important problems. I deny that it is all luck, but I admit there is a fair element of luck. I subscribe to Pasteur's "Luck favors the prepared mind." I favor heavily what I did. Friday afternoons for years — great thoughts only — means that I committed 10% of my time trying to understand the bigger problems in the field, i.e. what was and what was not important. I found in the early days I had believed “this” and yet had spent all week marching in “that” direction. It was kind of foolish. If I really believe the action is over there, why do I march in this direction? I either had to change my goal or change what I did. So I changed something I did and I marched in the direction I thought was important. It's that easy. Now you might tell me you haven't got control over what you have to work on. Well, when you first begin, you may not. But once you're moderately successful, there are more people asking for results than you can deliver and you have some power of choice, but not completely. I'll tell you a story about that, and it bears on the subject of educating your boss. I had a boss named Schelkunoff; he was, and still is, a very good friend of mine. Some military person came to me and demanded some answers by Friday. Well, I had already dedicated my computing resources to reducing data on the fly for a group of scientists; I was knee deep in short, small, important problems. This military person wanted me to solve his problem by the end of the day on Friday. I said, "No, I'll give it to you Monday. I can work on it over the weekend. I'm not going to do it now." He goes down to my boss, Schelkunoff, and Schelkunoff says, "You must run this for him; he's got to have it by Friday." I tell him, "Why do I?" He says, "You have to." I said, "Fine, Sergei, but you're sitting in your office Friday afternoon catching the late bus home to watch as this fellow walks out that door." I gave the military person the answers late Friday afternoon. I then went to Schelkunoff's office and sat down; as the man goes out I say, "You see Schelkunoff, this fellow has nothing under his arm; but I gave him the answers." On Monday morning Schelkunoff called him up and said, "Did you come in to work over the weekend?" I could hear, as it were, a pause as the fellow ran through his mind of what was going to happen; but he knew he would have had to sign in, and he'd better not say he had when he hadn't, so he said he hadn't. Ever after that Schelkunoff said, "You set your deadlines; you can change them." One lesson was sufficient to educate my boss as to why I didn't want to do big jobs that displaced exploratory research and why I was justified in not doing crash jobs which absorb all the research computing facilities. I wanted instead to use the facilities to compute a large number of small problems. Again, in the early days, I was limited in computing capacity and it was clear, in my area, that a "mathematician had no use for machines." But I needed more machine capacity. Every time I had to tell some scientist in some other area, "No I can't; I haven't the machine capacity," he complained. I said "Go tell your Vice President that Hamming needs more computing capacity." After a while I could see what was happening up there at the top; many people said to my Vice President, "Your man needs more computing capacity." I got it! I also did a second thing. When I loaned what little programming power we had to help in the early days of computing, I said, "We are not getting the recognition for our programmers that they deserve. When you publish a paper you will thank that programmer or you aren't getting any more help from me. That programmer is going to be thanked by name; she's worked hard." I waited a couple of years. I then went through a year of BSTJ articles and counted what fraction thanked some programmer. I took it into the boss and said, "That's the central role computing is playing in Bell Labs; if the BSTJ is important, that's how important computing is." He had to give in. You can educate your bosses. It's a hard job. In this talk I'm only viewing from the bottom up; I'm not viewing from the top down. But I am telling you how you can get what you want in spite of top management. You have to sell your ideas there also. Well I now come down to the topic, "Is the effort to be a great scientist worth it?" To answer this, you must ask people. When you get beyond their modesty, most people will say, "Yes, doing really first-class work, and knowing it, is as good as wine, women and song put together," or if it's a woman she says, "It is as good as wine, men and song put together." And if you look at the bosses, they tend to come back or ask for reports, trying to participate in those moments of discovery. They're always in the way. So evidently those who have done it, want to do it again. But it is a limited survey. I have never dared to go out and ask those who didn't do great work how they felt about the matter. It's a biased sample, but I still think it is worth the struggle. I think it is very definitely worth the struggle to try and do first-class work because the truth is, the value is in the struggle more than it is in the result. The struggle to make something of yourself seems to be worthwhile in itself. The success and fame are sort of dividends, in my opinion. I've told you how to do it. It is so easy, so why do so many people, with all their talents, fail? For example, my opinion, to this day, is that there are in the mathematics department at Bell Labs quite a few people far more able and far better endowed than I, but they didn't produce as much. Some of them did produce more than I did; Shannon produced more than I did, and some others produced a lot, but I was highly productive against a lot of other fellows who were better equipped. Why is it so? What happened to them? Why do so many of the people who have great promise, fail? Well, one of the reasons is drive and commitment. The people who do great work with less ability but who are committed to it, get more done that those who have great skill and dabble in it, who work during the day and go home and do other things and come back and work the next day. They don't have the deep commitment that is apparently necessary for really first-class work. They turn out lots of good work, but we were talking, remember, about first-class work. There is a difference. Good people, very talented people, almost always turn out good work. We're talking about the outstanding work, the type of work that gets the Nobel Prize and gets recognition. The second thing is, I think, the problem of personality defects. Now I'll cite a fellow whom I met out in Irvine. He had been the head of a computing center and he was temporarily on assignment as a special assistant to the president of the university. It was obvious he had a job with a great future. He took me into his office one time and showed me his method of getting letters done and how he took care of his correspondence. He pointed out how inefficient the secretary was. He kept all his letters stacked around there; he knew where everything was. And he would, on his word processor, get the letter out. He was bragging how marvelous it was and how he could get so much more work done without the secretary's interference. Well, behind his back, I talked to the secretary. The secretary said, "Of course I can't help him; I don't get his mail. He won't give me the stuff to log in; I don't know where he puts it on the floor. Of course I can't help him." So I went to him and said, "Look, if you adopt the present method and do what you can do single-handedly, you can go just that far and no farther than you can do single-handedly. If you will learn to work with the system, you can go as far as the system will support you." And, he never went any further. He had his personality defect of wanting total control and was not willing to recognize that you need the support of the system. You find this happening again and again; good scientists will fight the system rather than learn to work with the system and take advantage of all the system has to offer. It has a lot, if you learn how to use it. It takes patience, but you can learn how to use the system pretty well, and you can learn how to get around it. After all, if you want a decision “No”, you just go to your boss and get a “No” easy. If you want to do something, don't ask, do it. Present him with an accomplished fact. Don't give him a chance to tell you “No”. But if you want a “No”, it's easy to get a “No”. Another personality defect is ego assertion and I'll speak in this case of my own experience. I came from Los Alamos and in the early days I was using a machine in New York at 590 Madison Avenue where we merely rented time. I was still dressing in western clothes, big slash pockets, a bolo and all those things. I vaguely noticed that I was not getting as good service as other people. So I set out to measure. You came in and you waited for your turn; I felt I was not getting a fair deal. I said to myself, "Why? No Vice President at IBM said, ‘Give Hamming a bad time’. It is the secretaries at the bottom who are doing this. When a slot appears, they'll rush to find someone to slip in, but they go out and find somebody else. Now, why? I haven't mistreated them." Answer: I wasn't dressing the way they felt somebody in that situation should. It came down to just that — I wasn't dressing properly. I had to make the decision — was I going to assert my ego and dress the way I wanted to and have it steadily drain my effort from my professional life, or was I going to appear to conform better? I decided I would make an effort to appear to conform properly. The moment I did, I got much better service. And now, as an old colorful character, I get better service than other people. You should dress according to the expectations of the audience spoken to. If I am going to give an address at the MIT computer center, I dress with a bolo and an old corduroy jacket or something else. I know enough not to let my clothes, my appearance, my manners get in the way of what I care about. An enormous number of scientists feel they must assert their ego and do their thing their way. They have got to be able to do this, that, or the other thing, and they pay a steady price. John Tukey almost always dressed very casually. He would go into an important office and it would take a long time before the other fellow realized that this is a first-class man and he had better listen. For a long time John has had to overcome this kind of hostility. It's wasted effort! I didn't say you should conform; I said "The appearance of conforming gets you a long way." If you chose to assert your ego in any number of ways, "I am going to do it my way," you pay a small steady price throughout the whole of your professional career. And this, over a whole lifetime, adds up to an enormous amount of needless trouble. By taking the trouble to tell jokes to the secretaries and being a little friendly, I got superb secretarial help. For instance, one time for some idiot reason all the reproducing services at Murray Hill were tied up. Don't ask me how, but they were. I wanted something done. My secretary called up somebody at Holmdel, hopped the company car, made the hour-long trip down and got it reproduced, and then came back. It was a payoff for the times I had made an effort to cheer her up, tell her jokes and be friendly; it was that little extra work that later paid off for me. By realizing you have to use the system and studying how to get the system to do your work, you learn how to adapt the system to your desires. Or you can fight it steadily, as a small undeclared war, for the whole of your life. And I think John Tukey paid a terrible price needlessly. He was a genius anyhow, but I think it would have been far better, and far simpler, had he been willing to conform a little bit instead of ego asserting. He is going to dress the way he wants all of the time. It applies not only to dress but to a thousand other things; people will continue to fight the system. Not that you shouldn't occasionally! When they moved the library from the middle of Murray Hill to the far end, a friend of mine put in a request for a bicycle. Well, the organization was not dumb. They waited awhile and sent back a map of the grounds saying, "Will you please indicate on this map what paths you are going to take so we can get an insurance policy covering you." A few more weeks went by. They then asked, "Where are you going to store the bicycle and how will it be locked so we can do so and so." He finally realized that of course he was going to be red-taped to death so he gave in. He rose to be the President of Bell Laboratories. Barney Oliver was a good man. He wrote a letter one time to the IEEE. At that time the official shelf space at Bell Labs was so much and the height of the IEEE Proceedings at that time was larger; and since you couldn't change the size of the official shelf space he wrote this letter to the IEEE Publication person saying, since so many IEEE members were at Bell Labs and since the official space was so high the journal size should be changed. He sent it for his boss's signature. Back came a carbon with his signature, but he still doesn't know whether the original was sent or not. I am not saying you shouldn't make gestures of reform. I am saying that my study of able people is that they don't get themselves committed to that kind of warfare. They play it a little bit and drop it and get on with their work. Many a second-rate fellow gets caught up in some little twitting of the system, and carries it through to warfare. He expends his energy in a foolish project. Now you are going to tell me that somebody has to change the system. I agree; somebody's has to. Which do you want to be? The person who changes the system or the person who does first-class science? Which person is it that you want to be? Be clear, when you fight the system and struggle with it, what you are doing, how far to go out of amusement, and how much to waste your effort fighting the system. My advice is to let somebody else do it and you get on with becoming a first-class scientist. Very few of you have the ability to both reform the system and become a first-class scientist. On the other hand, we can't always give in. There are times when a certain amount of rebellion is sensible. I have observed almost all scientists enjoy a certain amount of twitting the system for the sheer love of it. What it comes down to basically is that you cannot be original in one area without having originality in others. Originality is being different. You can't be an original scientist without having some other original characteristics. But many a scientist has let his quirks in other places make him pay a far higher price than is necessary for the ego satisfaction he or she gets. I'm not against all ego assertion; I'm against some. Another fault is anger. Often a scientist becomes angry, and this is no way to handle things. Amusement, yes, anger, no. Anger is misdirected. You should follow and cooperate rather than struggle against the system all the time. Another thing you should look for is the positive side of things instead of the negative. I have already given you several examples, and there are many, many more; how, given the situation, by changing the way I looked at it, I converted what was apparently a defect to an asset. I'll give you another example. I am an egotistical person; there is no doubt about it. I knew that most people who took a sabbatical to write a book, didn't finish it on time. So before I left, I told all my friends that when I come back, that book was going to be done! Yes, I would have it done — I'd have been ashamed to come back without it! I used my ego to make myself behave the way I wanted to. I bragged about something so I'd have to perform. I found out many times, like a cornered rat in a real trap, I was surprisingly capable. I have found that it paid to say, “Oh yes, I'll get the answer for you Tuesday,” not having any idea how to do it. By Sunday night I was really hard thinking on how I was going to deliver by Tuesday. I often put my pride on the line and sometimes I failed, but as I said, like a cornered rat I'm surprised how often I did a good job. I think you need to learn to use yourself. I think you need to know how to convert a situation from one view to another which would increase the chance of success. Now self-delusion in humans is very, very common. There are innumerable ways of you changing a thing and kidding yourself and making it look some other way. When you ask, "Why didn't you do such and such," the person has a thousand alibis. If you look at the history of science, usually these days there are ten people right there ready, and we pay off for the person who is there first. The other nine fellows say, "Well, I had the idea but I didn't do it and so on and so on." There are so many alibis. Why weren't you first? Why didn't you do it right? Don't try an alibi. Don't try and kid yourself. You can tell other people all the alibis you want. I don't mind. But to yourself try to be honest. If you really want to be a first-class scientist you need to know yourself, your weaknesses, your strengths, and your bad faults, like my egotism. How can you convert a fault to an asset? How can you convert a situation where you haven't got enough manpower to move into a direction when that's exactly what you need to do? I say again that I have seen, as I studied the history, the successful scientist changed the viewpoint and what was a defect became an asset. In summary, I claim that some of the reasons why so many people who have greatness within their grasp don't succeed are: they don't work on important problems, they don't become emotionally involved, they don't try and change what is difficult to some other situation which is easily done but is still important, and they keep giving themselves alibis why they don't. They keep saying that it is a matter of luck. I've told you how easy it is; furthermore I've told you how to reform. Therefore, go forth and become great scientists! ## Questions and Answers A. G. Chynoweth: Well that was 50 minutes of concentrated wisdom and observations accumulated over a fantastic career; I lost track of all the observations that were striking home. Some of them are very very timely. One was the plea for more computer capacity; I was hearing nothing but that this morning from several people, over and over again. So that was right on the mark today even though here we are 20 – 30 years after when you were making similar remarks, Dick. I can think of all sorts of lessons that all of us can draw from your talk. And for one, as I walk around the halls in the future I hope I won't see as many closed doors in Bellcore. That was one observation I thought was very intriguing. Thank you very, very much indeed Dick; that was a wonderful recollection. I'll now open it up for questions. I'm sure there are many people who would like to take up on some of the points that Dick was making. Hamming: First let me respond to Alan Chynoweth about computing. I had computing in research and for 10 years I kept telling my management, “Get that !&@#% machine out of research. We are being forced to run problems all the time. We can't do research because were too busy operating and running the computing machines.” Finally the message got through. They were going to move computing out of research to someplace else. I was persona non grata to say the least and I was surprised that people didn't kick my shins because everybody was having their toy taken away from them. I went in to Ed David's office and said, “Look Ed, you've got to give your researchers a machine. If you give them a great big machine, we'll be back in the same trouble we were before, so busy keeping it going we can't think. Give them the smallest machine you can because they are very able people. They will learn how to do things on a small machine instead of mass computing.” As far as I'm concerned, that's how UNIX arose. We gave them a moderately small machine and they decided to make it do great things. They had to come up with a system to do it on. It is called UNIX! A. G. Chynoweth: I just have to pick up on that one. In our present environment, Dick, while we wrestle with some of the red tape attributed to, or required by, the regulators, there is one quote that one exasperated AVP came up with and I've used it over and over again. He growled that, "UNIX was never a deliverable!" Question: What about personal stress? Does that seem to make a difference? Hamming: Yes, it does. If you don't get emotionally involved, it doesn't. I had incipient ulcers most of the years that I was at Bell Labs. I have since gone off to the Naval Postgraduate School and laid back somewhat, and now my health is much better. But if you want to be a great scientist you're going to have to put up with stress. You can lead a nice life; you can be a nice guy or you can be a great scientist. But nice guys end last, is what Leo Durocher said. If you want to lead a nice happy life with a lot of recreation and everything else, you'll lead a nice life. Question: The remarks about having courage, no one could argue with; but those of us who have gray hairs or who are well established don't have to worry too much. But what I sense among the young people these days is a real concern over the risk taking in a highly competitive environment. Do you have any words of wisdom on this? Hamming: I'll quote Ed David more. Ed David was concerned about the general loss of nerve in our society. It does seem to me that we've gone through various periods. Coming out of the war, coming out of Los Alamos where we built the bomb, coming out of building the radars and so on, there came into the mathematics department, and the research area, a group of people with a lot of guts. They've just seen things done; they've just won a war which was fantastic. We had reasons for having courage and therefore we did a great deal. I can't arrange that situation to do it again. I cannot blame the present generation for not having it, but I agree with what you say; I just cannot attach blame to it. It doesn't seem to me they have the desire for greatness; they lack the courage to do it. But we had, because we were in a favorable circumstance to have it; we just came through a tremendously successful war. In the war we were looking very, very bad for a long while; it was a very desperate struggle as you well know. And our success, I think, gave us courage and self confidence; that's why you see, beginning in the late forties through the fifties, a tremendous productivity at the labs which was stimulated from the earlier times. Because many of us were earlier forced to learn other things — we were forced to learn the things we didn't want to learn, we were forced to have an open door — and then we could exploit those things we learned. It is true, and I can't do anything about it; I cannot blame the present generation either. It's just a fact. Question: Is there something management could or should do? Hamming: Management can do very little. If you want to talk about managing research, that's a totally different talk. I'd take another hour doing that. This talk is about how the individual gets very successful research done in spite of anything the management does or in spite of any other opposition. And how do you do it? Just as I observe people doing it. It's just that simple and that hard! Question: Is brainstorming a daily process? Hamming: Once that was a very popular thing, but it seems not to have paid off. For myself I find it desirable to talk to other people; but a session of brainstorming is seldom worthwhile. I do go in to strictly talk to somebody and say, "Look, I think there has to be something here. Here's what I think I see ..." and then begin talking back and forth. But you want to pick capable people. To use another analogy, you know the idea called the “critical mass.” If you have enough stuff you have critical mass. There is also the idea I used to call “sound absorbers”. When you get too many sound absorbers, you give out an idea and they merely say, "Yes, yes, yes." What you want to do is get that critical mass in action; "Yes, that reminds me of so and so," or, "Have you thought about that or this?" When you talk to other people, you want to get rid of those sound absorbers who are nice people but merely say, "Oh yes," and to find those who will stimulate you right back. For example, you couldn't talk to John Pierce without being stimulated very quickly. There were a group of other people I used to talk with. For example there was Ed Gilbert; I used to go down to his office regularly and ask him questions and listen and come back stimulated. I picked my people carefully with whom I did or whom I didn't brainstorm because the sound absorbers are a curse. They are just nice guys; they fill the whole space and they contribute nothing except they absorb ideas and the new ideas just die away instead of echoing on. Yes, I find it necessary to talk to people. I think people with closed doors fail to do this so they fail to get their ideas sharpened, such as "Did you ever notice something over here?" I never knew anything about it — I can go over and look. Somebody points the way. On my visit here, I have already found several books that I must read when I get home. I talk to people and ask questions when I think they can answer me and give me clues that I do not know about. I go out and look! Question: What kind of tradeoffs did you make in allocating your time for reading and writing and actually doing research? Hamming: I believed, in my early days, that you should spend at least as much time in the polish and presentation as you did in the original research. Now at least 50% of the time must go for the presentation. It's a big, big number. Question: How much effort should go into library work? Hamming: It depends upon the field. I will say this about it. There was a fellow at Bell Labs, a very, very, smart guy. He was always in the library; he read everything. If you wanted references, you went to him and he gave you all kinds of references. But in the middle of forming these theories, I formed a proposition: there would be no effect named after him in the long run. He is now retired from Bell Labs and is an Adjunct Professor. He was very valuable; I'm not questioning that. He wrote some very good Physical Review articles; but there's no effect named after him because he read too much. If you read all the time what other people have done you will think the way they thought. If you want to think new thoughts that are different, then do what a lot of creative people do — get the problem reasonably clear and then refuse to look at any answers until you've thought the problem through carefully how you would do it, how you could slightly change the problem to be the correct one. So yes, you need to keep up. You need to keep up more to find out what the problems are than to read to find the solutions. The reading is necessary to know what is going on and what is possible. But reading to get the solutions does not seem to be the way to do great research. So I'll give you two answers. You read; but it is not the amount, it is the way you read that counts. Question: How do you get your name attached to things? Hamming: By doing great work. I'll tell you the hamming window one. I had given Tukey a hard time, quite a few times, and I got a phone call from him from Princeton to me at Murray Hill. I knew that he was writing up power spectra and he asked me if I would mind if he called a certain window a "hamming window." And I said to him, "Come on, John; you know perfectly well I did only a small part of the work but you also did a lot." He said, "Yes, Hamming, but you contributed a lot of small things; you're entitled to some credit." So he called it the hamming window. Now, let me go on. I had twitted John frequently about true greatness. I said true greatness is when your name is like ampere, watt, and fourier — when it's spelled with a lower case letter. That's how the hamming window came about. Question: Dick, would you care to comment on the relative effectiveness between giving talks, writing papers, and writing books? Hamming: In the short-haul, papers are very important if you want to stimulate someone tomorrow. If you want to get recognition long-haul, it seems to me writing books is more contribution because most of us need orientation. In this day of practically infinite knowledge, we need orientation to find our way. Let me tell you what infinite knowledge is. Since from the time of Newton to now, we have come close to doubling knowledge every 17 years, more or less. And we cope with that, essentially, by specialization. In the next 340 years at that rate, there will be 20 doublings, i.e. a million, and there will be a million fields of specialty for every one field now. It isn't going to happen. The present growth of knowledge will choke itself off until we get different tools. I believe that books which try to digest, coordinate, get rid of the duplication, get rid of the less fruitful methods and present the underlying ideas clearly of what we know now, will be the things the future generations will value. Public talks are necessary; private talks are necessary; written papers are necessary. But I am inclined to believe that, in the long-haul, books which leave out what's not essential are more important than books which tell you everything because you don't want to know everything. I don't want to know that much about penguins is the usual reply. You just want to know the essence. Question: You mentioned the problem of the Nobel Prize and the subsequent notoriety of what was done to some of the careers. Isn't that kind of a much more broad problem of fame? What can one do? Hamming: Some things you could do are the following. Somewhere around every seven years make a significant, if not complete, shift in your field. Thus, I shifted from numerical analysis, to hardware, to software, and so on, periodically, because you tend to use up your ideas. When you go to a new field, you have to start over as a baby. You are no longer the big mukity muk and you can start back there and you can start planting those acorns which will become the giant oaks. Shannon, I believe, ruined himself. In fact when he left Bell Labs, I said, "That's the end of Shannon's scientific career." I received a lot of flak from my friends who said that Shannon was just as smart as ever. I said, "Yes, he'll be just as smart, but that's the end of his scientific career," and I truly believe it was. You have to change. You get tired after a while; you use up your originality in one field. You need to get something nearby. I'm not saying that you shift from music to theoretical physics to English literature; I mean within your field you should shift areas so that you don't go stale. You couldn't get away with forcing a change every seven years, but if you could, I would require a condition for doing research, being that you will change your field of research every seven years with a reasonable definition of what it means, or at the end of 10 years, management has the right to compel you to change. I would insist on a change because I'm serious. What happens to the old fellows is that they get a technique going; they keep on using it. They were marching in that direction which was right then, but the world changes. There's the new direction; but the old fellows are still marching in their former direction. You need to get into a new field to get new viewpoints, and before you use up all the old ones. You can do something about this, but it takes effort and energy. It takes courage to say, “Yes, I will give up my great reputation.” For example, when error correcting codes were well launched, having these theories, I said, "Hamming, you are going to quit reading papers in the field; you are going to ignore it completely; you are going to try and do something else other than coast on that." I deliberately refused to go on in that field. I wouldn't even read papers to try to force myself to have a chance to do something else. I managed myself, which is what I'm preaching in this whole talk. Knowing many of my own faults, I manage myself. I have a lot of faults, so I've got a lot of problems, i.e. a lot of possibilities of management. Question: Would you compare research and management? Hamming: If you want to be a great researcher, you won't make it being president of the company. If you want to be president of the company, that's another thing. I'm not against being president of the company. I just don't want to be. I think Ian Ross does a good job as President of Bell Labs. I'm not against it; but you have to be clear on what you want. Furthermore, when you're young, you may have picked wanting to be a great scientist, but as you live longer, you may change your mind. For instance, I went to my boss, Bode, one day and said, "Why did you ever become department head? Why didn't you just be a good scientist?" He said, "Hamming, I had a vision of what mathematics should be in Bell Laboratories. And I saw if that vision was going to be realized, I had to make it happen; I had to be department head." When your vision of what you want to do is what you can do single-handedly, then you should pursue it. The day your vision, what you think needs to be done, is bigger than what you can do single-handedly, then you have to move toward management. And the bigger the vision is, the farther in management you have to go. If you have a vision of what the whole laboratory should be, or the whole Bell System, you have to get there to make it happen. You can't make it happen from the bottom very easily. It depends upon what goals and what desires you have. And as they change in life, you have to be prepared to change. I chose to avoid management because I preferred to do what I could do single-handedly. But that's the choice that I made, and it is biased. Each person is entitled to their choice. Keep an open mind. But when you do choose a path, for heaven's sake be aware of what you have done and the choice you have made. Don't try to do both sides. Question: How important is one's own expectation or how important is it to be in a group or surrounded by people who expect great work from you? Hamming: At Bell Labs everyone expected good work from me — it was a big help. Everybody expects you to do a good job, so you do, if you've got pride. I think it's very valuable to have first-class people around. I sought out the best people. The moment that physics table lost the best people, I left. The moment I saw that the same was true of the chemistry table, I left. I tried to go with people who had great ability so I could learn from them and who would expect great results out of me. By deliberately managing myself, I think I did much better than laissez faire. Question: You, at the outset of your talk, minimized or played down luck; but you seemed also to gloss over the circumstances that got you to Los Alamos, that got you to Chicago, that got you to Bell Laboratories. Hamming: There was some luck. On the other hand I don't know the alternate branches. Until you can say that the other branches would not have been equally or more successful, I can't say. Is it luck the particular thing you do? For example, when I met Feynman at Los Alamos, I knew he was going to get a Nobel Prize. I didn't know what for. But I knew darn well he was going to do great work. No matter what directions came up in the future, this man would do great work. And sure enough, he did do great work. It isn't that you only do a little great work at this circumstance and that was luck, there are many opportunities sooner or later. There are a whole pail full of opportunities, of which, if you're in this situation, you seize one and you're great over there instead of over here. There is an element of luck, yes and no. Luck favors a prepared mind; luck favors a prepared person. It is not guaranteed; I don't guarantee success as being absolutely certain. I'd say luck changes the odds, but there is some definite control on the part of the individual. Go forth, then, and do great work! --- # altair-upset: The Evolution of UpSet plots in Altair How I turned a Jupyter notebook into a full-fledged Python package for UpSet plots [Altair](https://altair-viz.github.io) is the only plotting library that I've felt like loved me back. Maybe ggplot2 would love me back as well, but it's locked away in the tower of R. I went on a bit of a lark when I stumbled upon the [original upset-altair-notebook from HMS-DBMI](https://github.com/hms-dbmi/upset-altair-notebook). It was just a Jupyter Notebook, and only worked with Altair 4, which met none of my requirements. # The Journey The original notebook was great, but I needed something I could quickly pip install and use across projects. Plus, I needed to use Altair 5 for [marimo](https://marimo.io), I wanted to future-proof this tool for the community. Here's how it evolved: ```bash # The old way git clone https://github.com/hms-dbmi/upset-altair-notebook conda env create -f environment.yml conda activate upset-altair-env jupyter notebook ``` ```bash # The new way pip install altair-upset ``` # Making it easy to install First step was packaging. I'm a big fan of not reinventing the wheel, so I kept the core visualization logic but wrapped it in a proper Python package structure. This meant: ```python import altair_upset as au # Create UpSet plot with one clean function call chart = au.UpSetAltair( data=my_data, sets=["gene_set1", "gene_set2", "gene_set3"], title="Gene Set Intersections" ) ``` # Hitting Save Before the Boss Fight Here's where it gets interesting - I created a snapshot of the Altair 4 functionality before diving into the Altair 5 boss battle. Think of it like saving your game before a major fight - if something goes wrong, you can always roll back to a working state. Why? Because I've learned from enough bioinformatics murder mysteries that breaking changes in dependencies can be a nightmare. I tried to just port the function to Altair 5 and wound up with some really weird functionality that I couldn't make sense of. The secret weapon here was [syrupy](https://github.com/syrupy-project/syrupy) - it let me capture the exact state of my plots before making any changes. # Migrating to Altair 5 I think the [diff between the versions tells most of the story](https://github.com/edmundmiller/altair-upset/compare/0.1.1...0.2.0). It was mostly just swapping out a properties calls that got moved. ![Shared Mutations of COVID Variants UpSet Plot](https://raw.githubusercontent.com/edmundmiller/altair-upset/3ab3e4de21fcaf02dd0ea0211cc14d08238a689b/tests/__snapshots__/test_covid_mutations/test_covid_mutations_subset%5Bimage%5D.png) # Lessons Learned Along the Way This was my first rodeo with creating a Python package from scratch, and boy was it a journey. Remember how I said Altair was the only plotting library that loved me back? Well, the Python packaging ecosystem... First off, PyPI takes security more seriously than that one PI who makes you wear a lab coat just to look at a computer. I accidentally pushed v0.2.0 and then tried to yank it back - spoiler alert: you can't. It's like trying to take back an email after hitting send. That version now lives in infamy, a permanent reminder that "move fast and break things" doesn't fly with package registries. But the real game-changer? Discovering [`uv`](https://docs.astral.sh/uv/) and [`pyproject.toml`](https://packaging.python.org/en/latest/guides/writing-pyproject-toml/). After years of fighting with `setup.py` and virtual environments (and drowning in documentation that felt like it was written for people who already knew everything), this felt like finding the cheat codes. Finally, Python package management that doesn't feel like solving a Rubik's cube in the dark. ```shell # The old way python setup.py develop # pray it works pip install -e . # pray harder python -m venv venv # why do I need this again? # The new way uv sync ``` If you're just starting out with Python packaging, do yourself a favor - skip the history lesson and jump straight to `pyproject.toml`. Your future self will thank you. # What's Next? I'm keeping this project lean and focused. Want a feature? PRs welcome! Check it out on GitHub: [altair-upset](https://github.com/edmundmiller/altair-upset) --- # Migration from Biocontainers to Seqera Containers: Part 2 nf-core containers automation: how it'll all work behind the curtain import { Image } from 'astro:assets'; import { YouTube } from '@astro-community/astro-embed-youtube'; import Admonition from '../../components/admonition.astro'; # Introduction In nf-core, we've been excited to adopt [Wave](https://seqera.io/wave/) to automate software container builds, and have been looking for the right way to do it. With the announcement of [Seqera Containers](https://seqera.io/containers/) we felt it was the right time to put in the effort to migrate our containers to be built using Wave, using Seqera Containers to host the container images for our modules. You can read more about our motivation for this change in [Part 1 of this blog post](https://nf-co.re/blog/2024/seqera-containers-part-1). Here, in Part 2, we will dig into the technical details: how it all works behind the curtain. You don't need to know or understand any of this as an end-user of nf-core pipelines, or even as a contributor to nf-core modules, but we thought it would be interesting to share the details. It's mostly to serve as an architectural plan for the nf-core maintainers and infrastructure teams. - Module contributors edit `environment.yml` files to update software dependencies - Containers are automatically build for Docker + Singularity, `linux/amd64` + `linux/arm64` - Conda lock-files are saved for more reproducible and faster conda environments - Details are stored in the module's `meta.yml` - Pipelines auto-generate Nextflow config files when modules are updated - Pipeline usage remains basically unchanged # The end goal Before we dig into how the details of how the automation will work, let's summarise the end goal of this migration. ## Glossary - [`linux/amd64`](https://en.wikipedia.org/wiki/X86-64): Regular intel CPUs (aka `x86_64`) - [`linux/arm64`](https://en.wikipedia.org/wiki/AArch64): ARM CPUs (eg. AWS Graviton, aka `AArch64`). Not Apple Silicon. - [Apptainer](https://apptainer.org/): Alternative to Singularity, uses same image format - [Mamba](https://mamba.readthedocs.io): Alternative to Conda, uses same conda environment files - [Conda lock files](https://docs.conda.io/projects/conda/en/latest/user-guide/tasks/manage-environments.html#identical-conda-envs): Explicit lists of packages, used to recreate an environment exactly. ## Usage summary Pipeline users will see almost no change in current behaviour, but have several new configuration profiles available. | nf-core profile | Status | Use case | | ---------------------- | ------------------------------------------------------ | ----------------------------------------------------------------- | | `docker` | Unchanged | Docker images for `linux/amd64` | | `podman` | Unchanged | Docker images for `linux/amd64` | | `shifter` | Unchanged | Docker images for `linux/amd64` | | `charliecloud` | Unchanged | Docker images for `linux/amd64` | | `docker_arm` | New | Docker images for `linux/arm64` | | `podman_arm` | New | Docker images for `linux/arm64` | | `shifter_arm` | New | Docker images for `linux/arm64` | | `charliecloud_arm` | New | Docker images for `linux/arm64` | | `singularity` | Unchanged | Singularity images for `linux/amd64` | | `apptainer` | Updated | Singularity images for `linux/amd64` (not Docker, as previously) | | `singularity_arm` | New | Singularity images for `linux/arm64` | | `apptainer_arm` | New | Singularity images for `linux/arm64` | | `singularity_oras` | New | Singularity images for `linux/amd64` using the `oras://` protocol | | `apptainer_oras` | New | Singularity images for `linux/amd64` using the `oras://` protocol | | `singularity_oras_arm` | New | Singularity images for `linux/arm64` using the `oras://` protocol | | `apptainer_oras_arm` | New | Singularity images for `linux/arm64` using the `oras://` protocol | | `conda` | Updated | Conda lock files for `linux/amd64` | | `mamba` | Updated | Conda lock files for `linux/amd64`, using Mamba | | `conda_arm` | New | Conda lock files for `linux/arm64` | | `mamba_arm` | New | Conda lock files for `linux/arm64`, using Mamba | | `conda_env` | New | Conda with local `environment.yml` resolution | | `mamba_env` | New | Conda with local `environment.yml` resolution, using Mamba | ## Conda lock files Conda lock files were mentioned in [Part I of this blog post](/blog/2024/seqera-containers-part-1#exceptionally-reproducible). > These pin the exact dependency stack used by the build, not just the top-level primary tool being requested. > This effectively removes the need for conda to solve the build and also ships md5 hashes for every package. > This will greatly improve the reproducibility of the software environments for conda users and the reliability of Conda CI tests. They look something like this: ```yaml title="FastQC Conda lock file for linux/amd64" # micromamba env export --explicit # This file may be used to create an environment using: # $ conda create --name --file # platform: linux-64 @EXPLICIT https://conda.anaconda.org/conda-forge/linux-64/_libgcc_mutex-0.1-conda_forge.tar.bz2#d7c89558ba9fa0495403155b64376d81 https://conda.anaconda.org/conda-forge/linux-64/libgomp-13.2.0-h77fa898_7.conda#abf3fec87c2563697defa759dec3d639 https://conda.anaconda.org/conda-forge/linux-64/_openmp_mutex-4.5-2_gnu.tar.bz2#73aaf86a425cc6e73fcf236a5a46396d https://conda.anaconda.org/conda-forge/linux-64/libgcc-ng-13.2.0-h77fa898_7.conda#72ec1b1b04c4d15d4204ece1ecea5978 # .. and so on ``` ## Singularity: oras or https? Unfamiliar with `oras://`? Don't worry, it's relatively new in the field. It's a new protocol to reference container images, similar to `docker://` or `shub://`. It allows Singularity to interact with any OCI ([Open Container Initiative](https://opencontainers.org/)) compliant registry to pull images. Using `oras` has some advantages: - Singularity handles pulls in the process task, rather than in the Nextflow head job - This means less resource usage on the head node, and more parallelisation - Singularity can use authentication to pull from private registries (see [Singularity docs](https://docs.sylabs.io/guides/main/user-guide/cli/singularity_registry.html) for more information). However, there are some downsides: - Shared cache Nextflow options such as `$NXF_SINGULARITY_CACHEDIR` and `$NXF_SINGULARITY_LIBRARYDIR` are not used - Singularity must be installed when downloading images for offline use - `oras://` is only supported by recent versions of Singularity / Apptainer As such, we will continue to use `https` downloads for Singularity `SIF` images for now. However, we will start to provide new `-profile singularity_oras` profiles for anyone who would prefer to fetch images using the newer `oras` protocol. If you'd like to know more, check out the amazing [bytesize talk](https://nf-co.re/events/2024/bytesize_singularity_containers_hpc) by Marco Claudio De La Pierre ([@marcodelapierre](https://github.com/marcodelapierre/)) from June 2024: ## Modules All nf-core pipelines use a single container per process, and the majority of processes are encapsulated within shared modules in the [nf-core/modules](https://github.com/nf-core/modules) repository. As such, we must start with containers at the module level. For the latest discussion and progress on _bulk-updating_ existing nf-core modules, see GitHub issue [nf-core/modules#6698](https://github.com/nf-core/modules/issues/6698). ### Changes to `main.nf` With this switch, we simplify the `container` declaration, listing only the default container image: Docker, for `linux/amd64`. There will no longer be any string interpolation or logic within the container string. The `container` string is **never edited by hand** and is fully handled by the modules automation. With [the FastQC module](https://github.com/nf-core/modules/blob/f768b283dbd8fc79d0d92b0f68665d7bed94cabc/modules/nf-core/fastqc/main.nf#L6-L8) as an example: ```diff title="main.nf" process FASTQC { label 'process_medium' conda "${moduleDir}/environment.yml" + container "fastqc:0.12.1--5cfd0f3cb6760c42" // automatically generated - container "${ workflow.containerEngine == 'singularity' && !task.ext.singularity_pull_docker_container ? - 'https://depot.galaxyproject.org/singularity/fastqc:0.12.1--hdfd78af_0' : - 'biocontainers/fastqc:0.12.1--hdfd78af_0' }" input: tuple val(meta), path(reads) ``` We considered removing both `conda` and `container` declarations from the module `main.nf` file entirely. However, we see benefit in keeping these in this form because: - It's clearer to those exploring the code about what the module requires. - The container string is needed to tie the module to the pipeline config files (see [Building config files](#building-config-files) below) Removing the container logic from the string should be a big win for readability. ### Changes to `meta.yml` Through the magic of automation, we will append and then validate the following fields within the module's `meta.yml` file. Following the [FastQC example](https://github.com/nf-core/modules/blob/f768b283dbd8fc79d0d92b0f68665d7bed94cabc/modules/nf-core/fastqc/meta.yml) from above: ```yaml title="meta.yml" # ..existing meta.yml content above containers: docker: linux_amd64: name: community.wave.seqera.io/library/fastqc:0.12.1--5cfd0f3cb6760c42 build_id: 5cfd0f3cb6760c42_1 scan_id: 6fc310277b74 linux_arm64: name: community.wave.seqera.io/library/fastqc:0.12.1--d3caca66b4f3d3b0 build_id: d3caca66b4f3d3b0_1 scan_id: d9a1db848b9b singularity: linux_amd64: name: oras://community.wave.seqera.io/library/fastqc:0.12.1--0827550dd72a3745 https: https://community-cr-prod.seqera.io/docker/registry/v2/blobs/sha256/b2/b280a35770a70ed67008c1d6b6db118409bc3adbb3a98edcd55991189e5116f6/data build_id: 0827550dd72a3745_1 linux_arm64: name: oras://community.wave.seqera.io/library/fastqc:0.12.1--b2ccdee5305e5859 https: https://community-cr-prod.seqera.io/docker/registry/v2/blobs/sha256/76/76e744b425a6b4c7eb8f12e03fa15daf7054de36557d2f0c4eb53ad952f9b0e3/data build_id: b2ccdee5305e5859_1 conda: linux_amd64: lock file: https://wave.seqera.io/v1alpha1/builds/5cfd0f3cb6760c42_1/condalock linux_arm64: lock file: https://wave.seqera.io/v1alpha1/builds/d3caca66b4f3d3b0_1/condalock ``` All images are all built at the same time, avoiding Conda dependency drift. The build and scan IDs allow us to trace back to the build logs and security scans for these images. The Conda lock files are a new addition to the nf-core ecosystem and will help reproducibility for Conda users. These are generated during the Docker image build and are specific to architecture. The lock files can be accessed remotely via the Wave API, so we can treat them much in the same way that we treat remote container images. ## Pipelines Container information at module-level is great, but it's not enough. Nextflow doesn't know about module `meta.yml` files (they're an nf-core invention), so we need to tie these into the pipeline code where they will run. The heart of the solution is to auto-generate a config file for each software packaging type (Docker, Singularity, Conda) and platform (`linux/arch64` and `linux/arm64`). These will be created by the nf-core/tools CLI and never be edited by hand, so no manual merging will be required. They'll simply be regenerated and overwritten every time a version of a module is updated. Each config file will specify the `container` or `conda` directive for every process in the pipeline: ```groovy title="config/containers_docker_amd64.config" // AUTOGENERATED CONFIG BELOW THIS POINT - DO NOT EDIT process { withName: 'NF_PIPELINE:FASTQC' { container = 'fastqc:0.12.1--5cfd0f3cb6760c42' } } process { withName: 'NF_PIPELINE:MULTIQC' { container = 'multiqc:1.25--9968ff4994a2e2d7' } } process { withName: 'NF_PIPELINE:ANALYSIS_PLOTS' { container = 'express_click_pandas_plotly_typing:58d94b8a8e79e144' } } //.. and so on, for each process in the pipeline ``` Likewise, the conda config files will point to the lock files for each process: ```groovy title="config/conda_lock files_amd64.config" // AUTOGENERATED CONFIG BELOW THIS POINT - DO NOT EDIT process { withName: 'NF_PIPELINE:FASTQC' { conda = 'https://wave.seqera.io/v1alpha1/builds/5cfd0f3cb6760c42_1/condalock' } } process { withName: 'NF_PIPELINE:MULTIQC' { conda = 'https://wave.seqera.io/v1alpha1/builds/9968ff4994a2e2d7_1/condalock' } } process { withName: 'NF_PIPELINE:ANALYSIS_PLOTS' { conda = 'https://wave.seqera.io/v1alpha1/builds/58d94b8a8e79e144_1/condalock' } } //.. and so on, for each process in the pipeline ``` The main `nextflow.config` file will import these config files, depending on the [profile selected](#usage-summary) by the person running the pipeline. Singularity will have separate config files and associated `-profile`s for both `oras` and `https` containers, so that users can choose which to use. Local modules and any edge-case shared modules that cannot use the Seqera Containers automation will need the pipeline developer to hardcode container names and conda lock files manually. These can be added to the above config files as long as they remain above the comment line: ```groovy // AUTOGENERATED CONFIG BELOW THIS POINT - DO NOT EDIT ``` We're also taking this opportunity to update the `apptainer` and `mamba` profiles too, they will import the exact same config files as the `singularity` and `conda` profiles. Here's roughly how the `nextflow.config` file with the `-profile` config includes will look: Boilerplate code (eg. disabling other container engines) has been removed from this blog post code snippet for clarity. It will still be included in the pipelines. We may move this whole code block into it's own separate config file with `includeConfig` so that the main `nextflow.config` file is easier to read. ```groovy title="nextflow.config" // Set container for docker amd64 by default includeConfig 'config/containers_docker_amd64.config' profiles { docker { docker.enabled = true // Use the default config/containers_docker_amd64.config } docker_arm { includeConfig 'config/containers_docker_linux_arm64.config' docker.enabled = true } // podman, shifter, charliecloud the same as docker - also with _arm versions singularity { includeConfig 'config/containers_singularity_linux_amd64.config' singularity.enabled = true } singularity_arm { includeConfig 'config/containers_singularity_linux_arm64.config' singularity.enabled = true } singularity_oras { includeConfig 'config/containers_singularity_oras_linux_amd64.config' singularity.enabled = true } singularity_oras_arm { includeConfig 'config/containers_singularity_oras_linux_arm64.config' singularity.enabled = true } apptainer { includeConfig 'config/containers_singularity_linux_amd64.config' apptainer.enabled = true } apptainer_arm { includeConfig 'config/containers_singularity_linux_arm64.config' apptainer.enabled = true } apptainer_oras { includeConfig 'config/containers_singularity_oras_linux_amd64.config' apptainer.enabled = true } apptainer_oras_arm { includeConfig 'config/containers_singularity_oras_linux_arm64.config' apptainer.enabled = true } conda { includeConfig 'config/conda_lock files_amd64.config' conda.enabled = true } conda_arm { includeConfig 'config/conda_lock files_arm64.config' conda.enabled = true } conda_env { conda.enabled = true // Use the environment.yml file in the module main.nf } mamba { includeConfig 'config/conda_lock files_amd64.config' conda.enabled = true conda.useMamba = true } mamba_arm { includeConfig 'config/conda_lock files_arm64.config' conda.enabled = true conda.useMamba = true } mamba_env { conda.enabled = true // Use the environment.yml file in the module main.nf conda.useMamba = true } } docker.registry = 'community.wave.seqera.io/library' podman.registry = 'community.wave.seqera.io/library' apptainer.registry = 'oras://community.wave.seqera.io/library' singularity.registry = 'oras://community.wave.seqera.io/library' ``` Note that there are a few changes here: - New profiles with `_arm` suffixes for `linux/arm64` architectures - New profiles for `_oras` suffixes for using the `oras://` protocol - The `apptainer` profiles now uses the `singularity` config files - The `conda` profiles now use Conda lock files instead of `environment.yml` files - New `conda_env` profiles for those wanting to keep the old behaviour - New `mamba` profiles, using the `conda` config files - Base registries set to Seqera Containers Because we're only defining the image name and making use of the base container registry config option, it should still be simple to mirror containers to custom Docker registries and overwrite only `docker.registry` as before. # Automation - Modules The nf-core community loves automation. It's baked into the core of our community from our shared interest in automating workflows. We have linting bots, template updates, slack workflows, pipeline announcements. You name it, we've automated it. In these sections, we'll cover _how_ we're going to build all of these shiny new things without manual intervention. For the latest updates on modules container automation, see [nf-core/modules#6694](https://github.com/nf-core/modules/issues/6694). ## Updating conda packages The automation begins when a contributor wants to add a piece of software to a container. For instance, they decided that they need samtools installed. The contributor updates the `environment.yml` and adds a line with samtools: ```yml title="environment.yml" {6} channels: - conda-forge - bioconda dependencies: - bioconda::fastqc=0.12.1 - bioconda::samtools=1.16.1 ``` That will kick off the container image generation factory (it could equally be a change to remove a package, or change a pinned version). A commit will be pushed automatically with an updated `meta.yml` file pointing to the new containers, plus new nf-test snapshots for the software version checks. Only the interesting part needs to be edited by the developer (which tools to use) and all other steps are fully automated. ## Container image creation As with most automation in nf-core, container creation will happen in GitHub Actions. Edits to a module's `environment.yml` file will trigger a workflow that uses the [`wave-cli`](https://github.com/seqeralabs/wave-cli) to build the container images. 1. GitHub Actions identifies changes in the `environment.yml` file. 2. `wave-cli` is executed on the updated environment file. 3. Seqera Containers builds new containers for various platforms and architectures. 4. GitHub Actions runs stub tests commits the updated the [version snapshot](https://github.com/nf-core/modules/blob/1fe2e6de89778971df83632f16f388cf845836a9/modules/nf-core/bowtie/align/tests/main.nf.test.snap#L32-L46). ![Container creation flow](https://raw.githubusercontent.com/nf-core/website/main/sites/main-site/src/content/blog/2024/images/seqera-containers-part-2/creation_flow.excalidraw.svg) Once this GitHub Actions run completes it will push the new commits back to the PR and the regular nf-test CI will run. ## nf-test versions snapshot One of the primary reasons that we were so excited to adopt nf-test was the snapshot functionality. Every test has a snapshot file with the expected outputs, and the outputs are deterministic (not a binary file, and there's no dates). In that snapshot, we also capture the versions of the dependencies for the module ([example shown for bowtie2](https://github.com/nf-core/modules/blob/1fe2e6de89778971df83632f16f388cf845836a9/modules/nf-core/bowtie/align/tests/main.nf.test.snap#L32-L46)): ```json title="main.nf.test.snap" {4-7} "versions": { "content": [ { "BOWTIE_ALIGN": { "bowtie": "1.3.0", "samtools": "1.16.1" } } ], "meta": { "nf-test": "0.9.0", "nextflow": "24.04.4" }, "timestamp": "2024-09-27T10:42:58.892298" }, ``` This gives a second level of confirmation that the containers were correctly generated. When updating containers, the nf-test snapshot is parsed and compared to the snapshot from before the new containers were built. The snapshot changes are discarded if anything other than the `versions` key changed, so it should only vary in the software versions reported. If any other changes are detected, the snapshot change will be rejected and _not_ committed back to the PR. Then PR reviewers will see failing tests and need to manually update the snapshot file. This means that we can automatically commit the updated snapshot file in the PR if the tool output is unchanged, saving the developer from taking this extra step. Note that in the above example, the versions are in the snapshot in plain text, not using an md5 hash. This is a change that we will roll out for all modules, as it makes verification in the PR much easier. ## Automatic version bumps with Renovate We've recently adopted [Renovate](https://renovatebot.com/), a tool for automated dependency updates. It's multi-platform and multi-language and has become pretty popular in the devops space. It's similar to [GitHub's dependabot](https://docs.github.com/en/code-security/getting-started/dependabot-quickstart-guide#about-dependabot), but supports more languages and frameworks, and more importantly for nf-core, enables us to write our own custom dependencies. Renovate runs on a schedule and automatically updates software versions for us based on the specifications we've laid out in a [common config](https://github.com/nf-core/ops/blob/main/.github/renovate/default.json5). The magic starts with some nf-core automation to add renovate comments to the `environment.yml` file: ```yml {5,7,9} channels: - conda-forge - bioconda dependencies: # renovate: datasource=conda depName=bioconda/bwa - bioconda::bwa=0.7.18 # renovate: datasource=conda depName=bioconda/samtools - bioconda::samtools=1.20 # renovate: datasource=conda depName=bioconda/htslib - bioconda::htslib=1.20.0 ``` [These comments will be added](https://github.com/nf-core/modules/issues/6504) through [the batch module updates](https://github.com/nf-core/modules/issues/5828) happening this year. Future modules will have these comments [added automatically and linted](https://github.com/nf-core/tools/issues/3184) by the nf-core/tools CLI. The comments allow some scary regexes to find the conda dependencies and their versions in nf-core/modules, and check if there's a new version available. If there is a new version available, the Renovate bot will create a PR bumping the version, which in turn will kick off the container creation GitHub Action. The process will be very similar to the diagram laid out above, however we can go a step further: if the new software versions have no effect on the results of the tests, the PR will be automatically merged: ![Container renovation flow](https://raw.githubusercontent.com/nf-core/website/main/sites/main-site/src/content/blog/2024/images/seqera-containers-part-2/renovate_flow.excalidraw.svg) So: if all tests pass, the pull request is automatically merged without human intervention. In case of test failures, the Renovate bot automatically requests a review from the appropriate module maintainer using the `CODEOWNERS` file. The maintainer then steps in to fix failing tests and request a final review before merging. This efficient process ensures that software dependencies stay current with minimal manual oversight, reducing noise and streamlining development workflows. This will hopefully be the end of the _"can I get a review on this version bump"_ requests in `#review-requests`! # Automation - Pipelines We now have nice, up to date software packaging for the shared nf-core/modules with minimal manual intervention. However, we need to propagate these changes to the pipelines that use these modules. There are two main areas that we need to address: ## Building config files As [described above](#pipelines), pipelines will have a set of config files automatically generated that specify the container or conda environment for each process in the pipeline. Creation of these files will be triggered whenever installing, updating or removing a module, via the `nf-core` CLI. The config files will be completely regenerated each time, so there will never be any manual merging required. The trickiest part of this process is linking the module containers to the pipeline processes. Modules can be imported into pipelines with any alias or scope. We need to match this against the values that we find in the module `meta.yml` files: - Run `nextflow inspect` to generate default Docker `linux/arch64` config - Copy this file and replace the container names with the relevant container for each platform, using the `meta.yml` files that match the Docker container name This is why we duplicate the default docker container in both `main.nf` and `meta.yml` files - it allows us to link the module to the pipeline. Currently, `nextflow inspect` works by executing a dry-run of the pipeline. We can run using `-profile test`, but any processes that are skipped due to pipeline logic will not be included in the config file. This is a known limitation and we are currently working hard with the Nextflow team at Seqera on a new version of `nextflow inspect` which will return all processes, regardless of whether they are skipped or not. This will be a requirement for progression with the Seqera Containers migration. This process of copying files and using string substituion is a bit of a hack. If you have any ideas on how to improve this, please let us know! ## Edge cases and local modules There will always be edge cases that won't fit the automation described above. Not all software tools can be packaged on Bioconda (though we encourage it where possible!). For example, some tools have licensing restrictions that prevent them from being distributed in this way. Other edge-cases include optional usage of container variants for GPU support, or other hardware. Local modules won't be able to benefit from the Wave CLI automation to fetch containers from Seqera Containers and will have to be manually updated by the pipeline developer. For these reasons, we will still support custom `container` declarations in modules without use of Seqera Containers. It will be up to the module contributors to ensure that these are correctly specified and kept up to date manually. These can be specified in the `main.nf` file and added to the autogenerated platform-specific config files, as long as they remain above the comment line: ```groovy // AUTOGENERATED CONFIG BELOW THIS POINT - DO NOT EDIT ``` If at all possible then software should be packaged with Bioconda and Seqera Containers. Failing that, custom containers should be stored under the [nf-core account on quay.io](https://quay.io/organization/nf-core). The only time other docker registries / accounts should be used are if there are licensing issues restricting software redistribution. Custom containers can be also be built using Wave in continuous integration, it's just that they can't be pushed to the Seqera Containers registry. However, they _can_ be pushed to [quay.io](https://quay.io/organization/nf-core) automatically. We can do this using a similar mechanism to the automation used for changing `environment.yml` files, simply replacing it with `Dockerfile`s (see [nf-core/modules#4940](https://github.com/nf-core/modules/pull/4940)). ## Downloads The nf-core CLI has a `download` command that downloads pipeline code and software for offline use. This has a lot of hardcoded logic around the previous syntax of container strings and will need a significant rewrite. By the time we get to running this tool, the pipeline has all containers defined in configuration files. As such, we should be able to run the new and improved `nextflow inspect` command to return the container names for every process for a given platform. Once collected, we can download the images as before. The advantage of using this approach is that the download logic can be far simpler. Complex syntax for container strings is not a problem, as we can rely on Nextflow to resolve these to simple strings. # Roadmap This blog post lays out a future vision for this project. It serves as both a rubber-duck for the authors, a place to request feedback from the community, and as a roadmap for developers. There are many pieces of work that must come together for its completion, including but not limited to: {/* TODO: Am I missing any? */} - Nextflow - Improve `nextflow inspect` to return all processes - nf-core/tools - [Add and lint module Renovate comments](https://github.com/nf-core/tools/issues/3184) - [Lint for nf-test snapshots](https://github.com/nf-core/tools/issues/2504) - Write automation for creating pipeline container config files - [Rewrite `nf-core download`](https://github.com/nf-core/tools/issues/3179) - nf-core/modules - [Add renovate comments to environment.yml](https://github.com/nf-core/modules/issues/6504) - [Bulk update modules to use Seqera Containers](https://github.com/nf-core/modules/issues/6698) - [Build automation for fetching Seqera Containers](https://github.com/nf-core/modules/issues/6694) If all goes well, we hope to have the majority of this work completed by the end of 2024. --- # Migration from Biocontainers to Seqera Containers: Part 1 What Seqera Containers is and why we want to move to it. import { YouTube } from '@astro-community/astro-embed-youtube'; import { Image } from 'astro:assets'; import Admonition from '../../components/admonition.astro'; ## Introduction Dear nf-core community, the core team would like to inform you about some upcoming changes in how we would like to handle software containers... :package: Software containers are a fundamental part of modern bioinformatics workflows. Nextflow supports multiple container platforms, but the two most commonly used are Docker and Singularity. When using containers, all software requirements for a specific Nextflow process are wrapped up into an image and referenced within the pipeline code. The process tasks run in isolation in the host environment, and the end user doesn't need to worry about installing software dependencies for the pipeline. The software builds are locked in time, making results highly reproducible over many years. To ensure maximum compatibility across different environments, nf-core pipelines ship with the full range of options: conda environments as well as Docker and Singularity images. These are typically built using [Bioconda](https://bioconda.github.io/) and [BioContainers](https://biocontainers.pro/). The Bioconda and BioContainer projects have been invaluable to nf-core's success. We're hugely grateful to all contributors for their work, as well as to Anaconda, Quay.io and the Galaxy project for hosting these resources. ## Change is on the breeze {/* Photo by Simone Secci on Unsplash. */} Bioconda and BioContainers have worked very well for nf-core. However, we are now looking to migrate to a new system: Seqera Containers. The motivation comes down to a few key reasons: - Difficulties with BioContainers - [Mulled (multi-package) images](#mulled-multi-package-images) - [Slow Singularity image availability](#time-to-singularity-image) - [BioContainers API](#biocontainers-api) - [Reliability of hosting](#reliability-of-hosting) - New features we'd like - [Conda lock files](#exceptionally-reproducible) - Simplified developer workflow - [Better transparency](#trust-and-transparency) With the new Seqera Containers setup we should be able to address these problems, as well as provide the additional features. A detailed description of the proposed mechanism will come in part 2 of this blog post. Here are the key points: - Developers will _only_ need to edit the conda `environment.yml` (no process `container`) - Container images will be built by Wave and stored in Seqera Containers - Build logs and source files will be shown on the nf-core website module pages - Conda lock-files will be used for Conda users and CI tests - Software releases will be automatically bumped in modules, using [Renovate](https://renovatebot.com/) - There should be almost no change in how end-users run nf-core pipelines {''} This move has been in the works for over a year now. It started with discussions between nf-core maintainers and Seqera developers about what the community needed. Seqera's open source [Wave](https://seqera.io/wave/) container tool brings convenience, but long term storage of container images and stable container URIs were identified as hard requirements. In response, Seqera developed _Seqera Containers_, a free community resource with hosting infrastructure funded by AWS. This service was launched in spring 2024. The adoption of Seqera Containers in nf-core was initially raised in the nf-core steering group, followed by discussion in the core team and then maintainers team (see the [nf-core governance structure](/governance)). At each step, the discussion informed the planned infrastructure and setup for this change, as well as the development of Seqera Containers and Wave. Several steps are still needed to complete this migration: - nf-core/modules automation for fetching images and pinning conda-lock files and `meta.yml` references - nf-core/tools tooling for pipeline config file generation, with pipeline template update - Bulk update of nf-core/modules to use Seqera Containers - Update of nf-core pipelines to use the new modules If all goes to plan, we hope to have all modules updated by the end of 2024. This work will be tracked in GitHub issues as usual, starting with [nf-core/modules#5832](https://github.com/nf-core/modules/issues/5832). This issue has already had a substantial amount of discussion, much of which has informed this blog post. It's not too late to add your feedback! So if you have any questions or ideas, let us know on Github or Slack. Note that images from Seqera Containers can already be used in nf-core/modules, with the old syntax of using a `container` declaration (see [example](https://github.com/nf-core/modules/blob/b6b54f3929b0ba4a7c02a49191308ce1d8351f0d/modules/nf-core/bedtools/genomecov/main.nf#L6-L8)). ## BioContainers and Seqera Containers Nearly all nf-core modules are bundled with a conda `environment.yml`, listing software package dependencies from [Bioconda](https://bioconda.github.io/) (eg. [FastQC](https://github.com/nf-core/modules/blob/f768b283dbd8fc79d0d92b0f68665d7bed94cabc/modules/nf-core/fastqc/environment.yml)). The [BioContainers project](https://biocontainers.pro/) conveniently builds Docker and Singularity images for all tools on Bioconda automatically, hosting the images publicly on [quay.io](quay.io) and the Galaxy FTP servers, respectively. This means that the nf-core module developer can find the matching [BioContainer](https://biocontainers.pro/) images for their Bioconda package, and then add the image URLs into the module's `main.nf` script (see [FastQC example](https://github.com/nf-core/modules/blob/f768b283dbd8fc79d0d92b0f68665d7bed94cabc/modules/nf-core/fastqc/main.nf#L6-L8)). The nf-core pipeline user then pulls the container images from quay.io or the Galaxy FTP server when they run the pipeline. {/* */} [Seqera Containers](https://seqera.io/containers/) was launched in spring 2024 (see [Nextflow Summit talk](https://summit.nextflow.io/2024/boston/agenda/05-23--whats-new-in-the-nextflow/), [Seqera blog post](https://seqera.io/blog/introducing-seqera-pipelines-containers/), [AWS blog post + podcast](https://aws.amazon.com/blogs/hpc/announcing-seqera-containers-for-the-bioinformatics-community/), and the [Nextflow Channels podcast](https://nextflow.io/podcast/2024/ep38_seqera_pipelines_containers.html) for more information). Seqera Containers allows anyone to request a container image based on Conda or PyPI packages. The image is built on demand and then saved so that subsequent requests return the exact same container image files. The Seqera Containers service has a lot in common with BioContainers. Both generate Docker and Singularity images from Bioconda. The main difference is _when_ those builds happen. Seqera Containers is built on top of [Wave](https://seqera.io/wave/) - an open-source tool developed by Seqera for **on-demand** generation of containers. When a new tool or version is requested, Wave builds a container using Conda and returns it. In contrast, BioContainers runs a build when a new package is created on Bioconda. Building on-demand gives greater flexibility and scalability, especially for multi-tool containers. nf-core pipeline users won't need to interact with Wave directly. We simply plan to replace the mechanism for supplying the default container images specified in pipelines. It's important to distinguish the differences between Wave and Seqera Containers. Here's a brief comparison of the features:
BioContainers Wave Seqera Containers
Support Bioconda packages
Support all conda channels
Support PyPI (pip) packages
Docker + Singularity support
Linux aarch64 and arm64 ⏳ In progress
Multi-package containers ✅ Mulled
Conda lock files generated
Container build logs ❌ CI logs short lived
Docker container security scans ✅ quay.io ✅ Trivy ✅ Trivy
SBOM manifests (software bill of materials)
Long storage duration ✅ * ❌ 72 hours cache ✅ *
Pull delay for conda packages ✅ instant ❌ ~2-3 minutes for build on first request ✅ instant
Stable image URIs ❌ Single-run
Software required for end user Docker / Singularity Wave CLI / Nextflow, Docker / Singularity Docker / Singularity
Offline support with downloaded images
Guaranteed identical conda builds in future image pulls
\* Forever is a long time. Seqera is committing to keeping images for a minimum of 5 years from the time of their creation. BioContainers has no public policy.
### Mulled (multi-package) images
{/* Photo by Hannah Pemberton on Unsplash. */} BioContainers is primarily set up to have a 1:1 relationship with Bioconda. This is great for bioinformatics tools as an image is created every time a package is published. This works well until you need to use more than one tool in a single process, particularly when the tools are from elsewhere in the conda ecosystem. For example, many bioinformatics tools leverage [samtools](https://www.htslib.org/) to convert file formats or sort reads on the fly. Others may pipe output to compression tools like [pigz](https://zlib.net/pigz/). To resolve this, BioContainers has the concept of “mulled” images (as in [mulled wine](https://en.wikipedia.org/wiki/Mulled_wine)). Unfortunately, generating mulled containers is not trivial.
Luke Pembleton, Nextflow Ambassador, has a great blog post summarising the dark art of [Finding the right mulled biocontainer](https://lpembleton.rbind.io/posts/mulled-biocontainers/). Galaxy also has [documentation on the topic](https://docs.galaxyproject.org/en/master/admin/container_resolvers.html). In short, to request an image, you need to scroll to the bottom of a large [CSV file full of hashes](https://github.com/BioContainers/multi-package-containers/blob/master/combinations/hash.tsv) and add a line with your requested conda packages. Once edited, the images are built on CI. The container then becomes available after a review and merge. It used to be that you had to hunt through the GitHub action log to find the URI for your new container. However, [Moriz E. Beber](https://github.com/Midnighter) from the nf-core community has created a webpage to [generate the name of the mulled container](https://midnighter.github.io/mulled) which gently guides you through the process of finding the name of your container and how to update a container if you want to bump the software versions. Once the image is built and you know its address, the final step is to go to nf-core/modules, create a pull-request with the updated containers, and bump the versions in the conda `environment.yml`. This system works and, despite its complexity, is now a familiar process to many nf-core developers. However, it does present some problems. It's a highly manual process to update software versions and the container declarations - just fetching the software for a module often ends up taking longer than writing the entire module. There are also significant delays in waiting for the images to become available. All in all, it can be a frustrating experience and we see this in the volume of Slack messages asking for help. It seems likely that this puts some people off from contributing to nf-core. Migrating nf-core to use Seqera Containers will make provisioning multi-package containers easier. The module developer will only need to edit the module `environment.yml` file and everything else will be fully automated and made available almost immediately. Packages can be used minutes after their release on Bioconda and we will have automation in place to bump the Conda + Seqera Containers packages when a tool is updated. ### Time to Singularity image BioContainers builds both Docker and Singularity images. Docker images are pushed to quay.io and are available almost immediately, whereas Singularity images are pushed to the Galaxy FTP server in a nightly build, taking up to 24 hours to become available. This delay is frustrating for developers, as it breaks the flow of working on an update. Once a new version of a tool is released, they have to wait for the Singularity image to become available before they can add it and test the module. In practice, this often means that module updates get stuck in limbo for days until the developer has time to come back and check if the image is available. Singularity images are generated on the fly with Seqera Containers. Developers will be able update the nf-core module minutes after the Bioconda package is released. ### BioContainers API To simplify adding single-package BioContainer images to nf-core modules, the nf-core/tools package queries the BioContainer API for a given tool version and fetches the image URIs. This process is repeated during linting of nf-core modules and pipelines, to ensure that there is not an accidental mismatch between Conda and Docker / Singularity software versions. Unfortunately, the BioContainer API has had a history of being slow and unreliable, with a lot of down time. This has the knock-on effect of causing many CI tests to fail, which is frustrating and delays the process of pull-request merges. By migrating to Seqera Containers we will have similar automation when linting modules, but it will use the Seqera Wave API instead. This has been built to scale to very high volumes and has an extensive and robust back-end with a high degree of monitoring. Note that the Wave API will _not_ be used when people run nf-core pipelines, this is only developers running linting and GitHub Actions automation during development. ### Reliability of hosting BioContainer Docker images are hosted publicly on [quay.io](http://quay.io). This service is provided free of charge, however in recent years we have had some reliability issues. This becomes more noticeable as the community grows and the number of users trying to pull images increases. Because quay.io is a huge service for whom we are only a tiny player, we have no recourse when this happens and have to just wait until it becomes available again. Seqera serves nf-core as its primary community: they will respond immediately to any problems or needs. All Seqera Container images (Docker and Singularity) are hosted on custom infrastructure built by Seqera with hosting provided by AWS, so the whole stack is within reach. ## Key features {/* Photo by Behnam Norouzi on Unsplash. */} ### Exceptionally reproducible Seqera container images are hosted on long-term infrastructure and because their URLs will be hardcoded into pipeline configuration, they should be highly reproducible. Even if new builds of the same software are created in the future using Wave (for example due to an update in the base infrastructure used by Wave) then the old images are still pinned by the pipeline and will continue to be used, much like the BioContainer URIs used currently. As container URIs will be stored in the module-level `meta.yml` file, any pipelines using the same shared nf-core/module will also be using the exact same container images. We will further improve reproducibility at the conda level by adopting the use of conda lock-files. These pin the exact dependency stack used by the build, not just the top-level primary tool being requested. This effectively removes the need for conda to solve the build and also ships md5 hashes for every package. This will greatly improve the reproducibility of the software environments for conda users and the reliability of Conda CI tests. ```yaml # This file may be used to create an environment using: # $ conda create --name --file # platform: linux-64 @EXPLICIT https://conda.anaconda.org/conda-forge/linux-64/_libgcc_mutex-0.1-conda_forge.tar.bz2#d7c89558ba9fa0495403155b64376d81 https://conda.anaconda.org/conda-forge/linux-64/libgomp-14.1.0-h77fa898_1.conda#23c255b008c4f2ae008f81edcabaca89 https://conda.anaconda.org/conda-forge/linux-64/_openmp_mutex-4.5-2_gnu.tar.bz2#73aaf86a425cc6e73fcf236a5a46396d https://conda.anaconda.org/conda-forge/linux-64/libgcc-14.1.0-h77fa898_1.conda#002ef4463dd1e2b44a94a4ace468f5d2 # .. truncated .. ``` ### Trust and transparency Reproducibility is one thing, but it's also important that the images we use can be trusted and are secure. Users and developers must be able to verify that the contents of the container match the `environment.yml` for the module. The first point of trust is the `community.wave.seqera.io` base URI. This is only used for Seqera Containers and only `wave.seqera.io` is able to push images. It is also more locked-down than regular Wave, only allowing Conda builds and not custom Dockerfiles or augmented images. Second, the image URI includes a tag that has a hash of the input files used to request the image. We can use this to retrieve build details, which we will add to the nf-core website module page: - Build time, architecture and platform (Docker / Singularity) - `Dockerfile` / Singularity recipe and Conda `environment.yml` - Full build logs - Conda lock files with full package URLs and md5sums for entire dependency resolution - For Docker images: Trivy security scan and software bill of materials (SBOM) We haven't built this yet, but here's what the Wave build details page looks like, which has the same information: {/* Screenshot of a Wave container build details page. */} Wave build details page for the MultiQC v1.24.1 image: `community.wave.seqera.io/library/multiqc:1.24.1--789bc3917c8666da` The final part is the hash - from this we can infer the Wave Build ID and retrieve the build details: [Wave build details for 789bc3917c8666da_1](https://wave.seqera.io/view/builds/789bc3917c8666da_1) The new conda lock files also allow enhanced reproducibility for Conda users and more stable CI tests - we'll come back to these in a future blog post / bytesize. Finally, the image creation will happen automatically on GitHub Actions when an `environment.yml` file is edited. This means that the entire flow from conda file through to final containers will be entirely transparent, as the request to Wave itself and its response will be available in the GitHub Actions logs. ### No lock-in Using Seqera Containers does not mean that nf-core will be locked into using Seqera tooling. The suggested implementation has two components: 1. Automated on-demand generation of container images using Wave 2. Long term hosting images on Seqera Containers The tooling for image generation is [open-source](https://github.com/seqeralabs/wave). Seqera is hosting this API and build service for the community for free, but there's nothing to stop us from hosting the same API ourselves in the future. As such, there is no lock-in on this API or any tooling we will write around it. Long-term hosting is more fixed, but again Wave is designed to work with _any_ container registry. So if we want to change in the future we simply flip a configuration value and the generated images will be stored (“frozen”) at an alternative registry of our choosing. Old pipelines would have their configuration overwritten to use a different base registry. This is the same situation that we have currently with BioContainers and quay.io / Galaxy hosting. The nf-core community will remain free to pick and choose hosting solutions as needed. ### Special cases Not all tools can use Biocontainers, and we're aware of some special-case packages that have only bespoke Docker images. These will continue to work as they do today. For people who need to mirror container images to their own custom registry, this will still be possible. Changing the registry base will continue to work exactly the same way it does today. ## What this means for you ### Users If you're using nf-core pipelines but not developing them, then nothing should really change! All nf-core modules and pipelines should start supporting linux/arm64 CPU architectures, such as AWS Graviton / Raspberry Pis(!). Conda users should benefit from faster environment resolution, with more reproducible and stable software thanks to use of the lock files. Finally, the failure rate of docker pulls will improve as we phase out our usage of quay.io. ### Developers If you're a maintainer of a pipeline that uses nf-core/modules, this means that your life is about to get easier! Never again will you need to try to figure out how to make a mulled image, or wonder when your Singularity image build will be available. New software releases can be used with modules within minutes of release and most updates should happen automatically - freeing up maintainers from having to deal with routine package updates and reviews. In part-two of this blog post we will dive into the details of how we intend to build this tooling infrastructure in nf-core, so please check that out if you're interested. ## In conclusion We hope that the nf-core community is excited about these proposals, if you have any questions or concerns then please let us know in the nf-core Slack ❤️ {/* Photo by Providence Doucet on Unsplash. */} --- # Pixi for Bioinformatics: Bioconda on HPC and Local Machines Set up Pixi with conda-forge and Bioconda on an HPC cluster or laptop, install Nextflow, define tasks, and import an existing environment.yml. Recently, Anaconda has decided to start charging users for access to the default channel. This change in policy has significant implications for the data science and machine learning community, as Anaconda is a widely used platform for managing Python and R packages.[^1][^2][^3] Prefix.dev has a [nice article summarizing the whole thing](https://prefix.dev/blog/towards_a_vendor_lock_in_free_conda_experience). The Python ecosystem has recently seen a proliferation of attempts to completely rework its package managers. Several new projects have emerged, each aiming to address in the existing package managers. An issue that isn't really due that "pip and conda are bad". `pip` was released **April 4th, 2011** and `Anaconda` was first release **July 17th, 2012**. That's long before I even took an intro CS course my freshman year of college, so I'm not gonna judge what everyone else was doing then. The space is ripe for disruption. There have been some exciting complete rewrites that have come out recently with [ruff](https://docs.astral.sh/ruff) and [uv](https://docs.astral.sh/uv/) from Astral in the Python ecosystem. In the conda ecosystem, I've lost track of the different ways to install a conda package, Anaconda, miniconda, Mamba, micromamba, mambaforge. I can't keep up. I recently started using [`Pixi`](https://pixi.sh/latest/) from [prefix.dev](https://prefix.dev/). It's been really nice. I want to forget about a package manager. Pixi let's me do that. # Pixi for Bioinformatics ## Set up on a server/locally ```bash curl -fsSL https://pixi.sh/install.sh | bash # source ~/.bashrc pixi config append default-channels conda-forge --global pixi config append default-channels bioconda --global pixi global install -c bioconda nextflow pixi g i rclone ``` Just thought I'd add a quick TL;DR here. This is how easy it is to get started. Works on your server, locally on your laptop. Makes me think of some other great [bioinformatics tools](https://www.nextflow.io/)... ## Things every Bioinformatician loves to see ### One way to install it just about everywhere Problem: Which way should you install conda? - Anaconda - miniconda - Mamba - Micromamba - mambaforge Solution: ```bash curl -fsSL https://pixi.sh/install.sh | bash ``` ### Environment activation is automatic Problem: In every new terminal users have to run `conda activate applied-genomics`.[^5] Solution: Every project has the same commands - `pixi shell` - `pixi shell-hook` - `pixi run` ### Tasks are built-in Problem: `snakemake --cores 4` Solution: ```toml [tasks] run = "snakemake --cores 4" upload = { cmd = "rclone sync results/ box:THK_LAB_DATA/results/", depends-on = ["run"] } ``` ```sh pixi run upload ``` Sure, you could have a Makefile, or reach for [just](https://github.com/casey/just). But that's just one more tool everyone you collaborate with has to learn. ### Storing environment details is taken care of for you Problem: `conda install` doesn't update the `environment.yml`[^4] Solution: `pixi add` updates the `pixi.toml` for you. But wait there's more! It locks the version to avoid major version bumps automatically! ```toml [dependencies] python = ">=3.12.5,<4" ``` In case python4 sneaks up on you. ### No built in lock file Problem: Versions in the `environment.yml` are only half the battle...[^6] Solution: `pixi add` updates the `pixi.lock` file as well behind the scenes! ## It's gotta be difficult to migrate to, right? With [pixi you can import `environment.yml` files into a pixi project](https://pixi.sh/latest/switching_from/conda/#automated-switching). ```bash pixi init --import environment.yml ``` This will create a new project with the dependencies from the `environment.yml` file. ## Related next steps If you're thinking about Pixi because you want more reproducible bioinformatics workflows, you may also want to: [^1]: https://x.com/NM_Reid/status/1825997577151525338 [^2]: https://www.theregister.com/2024/08/08/anaconda_puts_the_squeeze_on/ [^3]: https://www.linkedin.com/feed/update/urn:li:share:7229549722070310912/ [^4]: Which makes it really hard to reproduce research code as we've learned firsthand... [^5]: Remembering the environment name in every project is really the issue. [^6]: Going to avoid going off into the weeds here, there is [conda-lock](https://github.com/conda/conda-lock). Reach out [Jonathan Manning](https://github.com/pinin4fjords) for his upcoming TED talk on conda lock files. --- # Format Snakemake in Emacs with Apheleia and snakefmt Configure Apheleia in Doom Emacs to format Snakemake files with snakefmt, including a Nix package and the exact set-formatter! setup. Dooms Emacs recently switched to using [Apheleia](https://github.com/radian-software/apheleia) as it's default formatter, thanks to a huge effort from [Ellis Kenyő](https://elken.dev) to refactor the [format module](https://docs.doomemacs.org/latest/?#/modules/editor/format). I've been writing a bit of [Snakemake](https://snakemake.github.io/) for the [Applied Genomics Course](https://applied-genomics.dev/) that I teach in the Summer. The last time that I wrote much Snakemake(circa 2020), [snakefmt](https://github.com/snakemake/snakefmt) didn't exist yet. It follows the design and specifications of Black. ## Aside: Packaging up `snakefmt` with Nix I've chosen the path less traveled, and use NixOS. Meaning I've sworn off using Conda, global pip installs, and other small conviences in the already narrow path that is using linux as a desktop environment. While Snakemake itself is packaged up in nixpkgs, snakefmt hasn't made it to the most magical repo on GitHub yet. Recently though I found a handy tool for quickly generating package derivations, [nix-init](https://github.com/nix-community/nix-init). ```sh nix run github:nix-community/nix-init -- --url https://github.com/snakemake/snakefmt ``` It pulls in the version and the dependencies, ✨Automagically✨. ```nix title="snakefmt.nix" { lib, python3, fetchFromGitHub, }: python3.pkgs.buildPythonApplication rec { pname = "snakefmt"; version = "0.10.2"; pyproject = true; src = fetchFromGitHub { owner = "snakemake"; repo = "snakefmt"; rev = "v${version}"; hash = "sha256-Sp48yedUiL8NCF7WF9QdvaOGocPXIBZ5bXXj7r4RVIM="; }; nativeBuildInputs = [ python3.pkgs.poetry-core ]; propagatedBuildInputs = with python3.pkgs; [ black click importlib-metadata toml ]; pythonImportsCheck = ["snakefmt"]; meta = with lib; { description = "The uncompromising Snakemake code formatter"; homepage = "https://github.com/snakemake/snakefmt"; changelog = "https://github.com/snakemake/snakefmt/blob/${src.rev}/CHANGELOG.md"; license = licenses.mit; maintainers = with maintainers; [edmundmiller]; mainProgram = "snakefmt"; }; } ``` ## Apheleia There are tons of code formatters out there. There's usually multiple for popular languages. Everyone's got an opinion on what style to use. Apheleia aims to remove specific Emacs packages for formatters. It's goal is to have one interface to run all of your formatters from Emacs. > running a code formatter on save suffers from the following two problems: > 1. It takes some time (e.g. around 200ms for Black on an empty file), which makes the editor feel less responsive. > 2. It invariably moves your cursor (point) somewhere unexpected if the changes made by the code formatter are too close to point's position. > Apheleia is an Emacs package which solves both of these problems comprehensively for all languages, allowing you to say goodbye to language-specific packages such as Blacken and prettier-js. The main opinion everyone shares is that a good code formatter should be fast, and therefore you should be able to forget about it. It's really simple to configure formatters in [Doom Emacs](https://github.com/doomemacs/doomemacs). ```emacs-lisp title="config.el" (set-formatter! 'snakefmt '("snakefmt" "-") :modes '(snakemake-mode)) ``` The `set-formatter!` macro takes: 1. The name you want to give the formatter `snakefmt` 2. The command you want ran (`--quiet` is used here often) `snakefmt -` 3. The modes you want associated with the formatter `snakemake-mode` Another example for [Alejandra](https://github.com/kamadorueda/alejandra) for Nix. ```emacs-lisp title="config.el" ;;; :lang nix (set-formatter! 'alejandra '("alejandra" "--quiet") :modes '(nix-mode)) ``` --- # Nextflow Emacs Workflow How I hack on Nextflow scripts using Emacs I was chatting with [David](https://dcgemperline.github.io/) at the recent Nextflow Summit in Boston about Emacs and comparing our various workflows. I thought I might turn this into a blog post after a question about [Emacs workflows on the Seqera discourse](https://community.seqera.io/t/what-is-your-nextflow-emacs-setup). When working with Nextflow and Emacs, my typical workflow involves having a local Emacs instance open with my project. I make changes to the source code within Emacs and run `nextflow run . -profile test,docker ...` in a separate terminal. I've tried integrated terminals in Emacs but it's just always _slightly off_, so I don't bother. Typically just have workspaces open with a Browser in workspace 1, Emacs in workspace 2, and a terminal open in workspace 3. For projects that require a remote system, such as an HPC cluster, I utilize [emacs-ssh-deploy](https://github.com/cjohansson/emacs-ssh-deploy). This tool automatically copies the file from my local machine to the appropriate directory on the server upon saving using `sftp`. To execute Nextflow, I simply ssh into the remote system and run `nextflow run . -profile ...` in a terminal. I have considered developing a [Transient](https://magit.vc/manual/transient/)-based launcher for [`nextflow-mode`](https://github.com/edmundmiller/nextflow-mode), similar to the one added to `snakemake-mode`. However, I have found these launchers to be fragile when dealing with terminal output. [I even created one for the NodeJS framework, Jest](https://github.com/edmundmiller/emacs-jest), but have not used it extensively myself. My goal is to keep my workflow simple and flexible, allowing for easy substitution of Emacs with other editors like Neovim, Helix, or Zed. As for my Emacs setup, I have been using [Doom Emacs](https://github.com/doomemacs/doomemacs) since 2018 [^1], and [my .doom.d configuration is available on GitHub](https://github.com/edmundmiller/.doom.d). The Nextflow specific parts of my config: ```elisp title="packages.el" (package! nextflow-mode :recipe (:host github :repo "edmundmiller/nextflow-mode")) ``` ```elisp title="config.el" (use-package! nextflow-mode :config (set-docsets! 'nextflow-mode "Groovy")) ``` [^1]: [Henrik](https://henrik.io/) really inspired my love of modules, and most of my development workflow. --- # GoatCounter vs Umami for Personal Site Analytics I compared GoatCounter, Umami, Plausible, and Fathom to find personal-site analytics without creepy tracking or another yak-shaving project. If you've ever experienced the itch to refresh your online presence, you'll know the feeling all too well. Recently, I gave my personal website a facelift (keep your eyes peeled for that upcoming blog post!), and took on the task of revamping [my partner's site](https://monimiller.com/) as well. During this period of digital rejuvenation, I stumbled upon a nugget of wisdom from the [Blogging for Devs](https://bloggingfordevs.com/) email course. They insist on setting up site analytics, not just for the sake of data, but for a source of motivation to keep pushing forward. This advice couldn't have come at a better time. Amidst the redesign, I found myself pondering over website analytics options. I have been utilizing [GoatCounter](https://www.goatcounter.com) for a while now, but as someone who strives to maintain efficiency (which I like to metaphorically describe as 'keeping my yak herd well shaven'), I'm always on the lookout for tools that streamline my workflow and bolster my motivation. Thus began my quest to find the perfect analytics tool to accompany my freshly updated sites. ## Requirements The list I'm looking to tick-off: 1. Not creepy 2. Self-hostable 3. But has a cloud option (I'm trying this new thing where I outsource my yak shaving) Martin Tournoij, the author of GoatCounter, has a [good piece on analytics on personal websites](https://www.arp242.net/personal-analytics.html). ## Initial Research I tried [Fathom](https://usefathom.com/ref/DYRELW) first. It seemed popular, and it was alright but the pricing at $15 a month was unreasonable for our web traffic currently. I'd tested out [Plausible](https://plausible.io) in the past. It looks cool, but there's no free tier. It's self-hostable, built on [ClickHouse](https://clickhouse.com/), it's flashy, but probably overkill. I was pretty excited to try [Umami](https://umami.is), the last time I had looked at it a while ago before it had a cloud hosting option. I think this one might be the only one to give GoatCounter a run for it's money. It has a hobby pricing of _free_ so it'll be hard to justify Plausible's $9 a month for the same 10K views. My first stop was [Fathom](https://usefathom.com/ref/DYRELW). It's garnered quite a following, perhaps for its ease of use and solid feature set. However, at $15 a month, I couldn't quite make peace with their pricing, especially given our modest web traffic. Then, my attention turned to [Plausible](https://plausible.io). There's no doubt it's slick, with a contemporary look and feel that can be quite enticing. Plus, it's built on the foundation of ClickHouse. However, the catch? They don't offer a free tier, and for something with such flash, I had to ask myself if it was too much for our needs. Finally, there was [Umami](https://umami.is). I had my eye on it some time ago, but back then, cloud hosting wasn't on the menu. What's especially exciting is its 'hobby' tier priced at _free_. When comparing that to Plausible's $9 per month for an equivalent 10K views, it's challenging to justify the cost of Plausible. ## Conclusion In summary, Umami is emerging as a strong contender. I'm going to use Plausible's free month to test Plausible, Umami, and GoatCounter together for a week or two since I've got it set up. Sorry for the slight performance hit for anyone reading this! 😬 ## **Update** Figured out GoatCounter added the ability to have multiple website under one login. This fixed my only grip with GoatCounter From the settings page: > Sites > Add GoatCounter to multiple websites by creating new sites. All sites will share the same users, and logins, but are otherwise completely separate. The current site’s settings are copied on creation, but are independent afterwards. > You can add as many as you want. --- # Using MDX in Doom Emacs Yo Dawg I heard you like Components so I put Components in your Markdown Following up [my recent post on using Astro in Emacs](./emacs-astro.md), I was missing a key feature. I needed better [MDX](https://mdxjs.com/) support. [Astro /heavily/ uses MDX](https://docs.astro.build/en/guides/markdown-content). ![DragonBall Meme](https://kagi.com/proxy/mdx-fusion-meme.jpg?c=Iyef1Nrg8olHNVNykoDod8U-LxFBgbddACMxorqhFpOgTCEGBGcZzgBkRZ23F9YB2D0tcBfnc8y-7Fp-y9-Fwly7uYBMnPCYOOSWYECNirDJxm2Rhyu3Z2oXo5pgR70NzkWPxUsibr4Ca0Tgt5qez9mSQByXfRUsdqVLM_hUT-8%3D) Following an article on [Configuring Emacs for MDX files](https://tailscale.dev/blog/configuring-emacs-mdx) which got the syntax highlighting there. --- # Setting up Doom Emacs for Astro Development Set up Astro in Doom Emacs with astro-ts-mode, Tree-sitter, lsp-mode, Apheleia and Prettier, plus Tailwind CSS IntelliSense. [Astro](https://astro.build/) is the new hot new web framework on the block. All the cool kids are using it. I've recently given up, drank the Kool-Aid, and gone all in on it. I've rewritten this website, [my partner's website](https://monimiller.com/), my university rugby club's website. I'm moving my _Applied Genomics_ course website to [Starlight](https://starlight.astro.build/), the Astro team's documentation framework. The [nf-core site](https://github.com/nf-core/website) has been rewritten in Astro and Svelte from PHP. _I'm all in_. The beauty of Astro is it's the [Nextflow](https://www.nextflow.io) of web frameworks.[^1] It allows you to wrap other UI Frameworks in a web framework rather than forcing you to pick one so you don't just have to pick React, Vue, or Svelte. You can have them all in the same application. You can just use [HTML components](https://docs.astro.build/en/basics/astro-components/#html-components). That's the beauty. That's why it's exciting. That's why I think it'll stick around. So anyways, I wanted to hook up Emacs with Astro support. For now, I've just been roughing it out there and running [Prettier](https://prettier.io/) by itself and turning off save on format and auto-complete. It's been scary. What I'm seeking from Emacs is multifaceted: Tree-sitter support, LSP (Language Server Protocol) support—to alert me of any missteps—and a fully functional formatter. A frustrating hour was lost to Prettier mangling my Astro templates by wrapping them in quotes—a bug I could have done without. And while we're at it, add Tailwind CSS LSP support into the mix for good measure. [^1]: Did I really just compare a very niche DSL to describe a niche programming language? ## Astro Tree-sitter Support > Tree-sitter is an incremental parsing system for programming tools. > Find out more about it on the [project's website](https://tree-sitter.github.io/tree-sitter/)! As the old saying goes, there's an Emacs package for everything. So, of course, someone's already written one for Astro and Tree-sitter. Setup for `astro-ts-mode` appears simple: ```elisp title="packages.el" (package! astro-ts-mode) ``` ```elisp title="config.el" (use-package! astro-ts-mode :after treesit-auto) ``` But wait there's [more](https://github.com/Sorixelle/astro-ts-mode?tab=readme-ov-file#setup)! > Because this major mode is powered by Tree-sitter, it depends on an external grammar to provide a syntax tree for Astro templates. To set it up, you’ll need to set treesit-language-source-alist to point to the correct repositories for each language. You can choose to set it up and run `treesit-install-language-grammar` for astro tsx and css. Or you can take the red pill and use [treesit-auto](https://github.com/renzmann/treesit-auto) and automatically install the language grammar. In case this your first Edmund experience, two things you should know. I love to automate things and I love a good rabbit hole. ### treesit-auto There was [a tip](https://github.com/Sorixelle/astro-ts-mode/issues/5) from [Ian S. Pringle](https://github.com/ispringle)(Who owns both a farm and a digital garden!). [Ruby Juric](https://github.com/Sorixelle/astro-ts-mode)(the author of `astro-ts-mode`) converted the snippet to use `let` to avoid creating a global variable. ```elisp title=config.el (use-package! astro-ts-mode :config (global-treesit-auto-mode) (let ((astro-recipe (make-treesit-auto-recipe :lang 'astro :ts-mode 'astro-ts-mode :url "https://github.com/virchau13/tree-sitter-astro" :revision "master" :source-dir "src"))) (add-to-list 'treesit-auto-recipe-list astro-recipe))) ``` I think this worked for me. I had built it manually with `treesit-auto` before. Oh by the way > !NOTE > Make sure you have a working C compiler as cc in your PATH, since this needs to compile the grammars. ## Emacs lsp-mode and Astro Language-server The official Astro editor docs link to [an article](https://medium.com/@jrmjrm/configuring-emacs-and-eglot-to-work-with-astro-language-server-9408eb709ab0) with instructions to configure eglot, but there's no equivalent one for lsp-mode. ```bash npm i -g @astrojs/language-server ``` I just had to add a hook to `astro-ts-mode` and it pulled right up. ```elisp title=config.el (use-package! astro-ts-mode :after treesit-auto :init (when (modulep! +lsp) (add-hook 'astro-ts-mode-hook #'lsp! 'append)) :config ;; ... ``` ## `prettier-plugin-astro` in Emacs with Apheleia From [Sorixelle's Emacs config](https://github.com/Sorixelle/dotfiles/blob/main/config/emacs-config.org#astro) I found the magic snippet that had prettier use `--parser=astro` in `.astro` files. ✨ ```elisp title=config.el (set-formatter! 'prettier-astro '("npx" "prettier" "--parser=astro" (apheleia-formatters-indent "--use-tabs" "--tab-width" 'astro-ts-mode-indent-offset)) :modes '(astro-ts-mode)) ``` ## Tailwind CSS IntelliSense in Emacs Of course there's already a package for [TailwindCSS using LSP](https://github.com/merrickluo/lsp-tailwindcss). With Doom Emacs installation instructions as well! ```elisp title="config.el" {"1. Launch the LSP in add-on-mode":4-5} {"2. Launch lsp-tailwindcss in astro-ts-mode":7-8} (use-package! lsp-tailwindcss :when (modulep! +lsp) :init (setq! lsp-tailwindcss-add-on-mode t) :config (add-to-list 'lsp-tailwindcss-major-modes 'astro-ts-mode)) ``` ## Conclusion You can find [all of the code in my Doom Emacs config](https://github.com/edmundmiller/.doom.d/tree/main/modules/lang/astro). It's got everything, Tree-sitter, LSP, Prettier, and Tailwind CSS IntelliSense. ## Related next steps If the Emacs and Astro crossover is what brought you here, these are good follow-ups: --- # Encrypt org-journal with age in Doom Emacs Set up age.el and rage in Doom Emacs, configure SSH key paths, and save org-journal entries as encrypted .org.age files when EasyPG hangs. # The problem Because [gpg 2.4.1 borked](https://dev.gnupg.org/T6481) Emacs\'s EasyPG. It just hangs on saving. ## Why not just use gpg? There are some ways around it, and I was using the `fset` hack until I read this [post about the person who corrupted their encrypted files](https://www.reddit.com/r/emacs/comments/18d6fmt/how_to_lock_yourself_out_of_a_gpg_encrypted_file/). I also had to run the elisp on every new Emacs instance. And then `lib-gcrypt` is marked as broken in NixOS if you use the `gnupg22` package(Version: 2.2.41), and a blowing past that stop sign sounded like a bad idea. So I started thinking outside the box. ## Why [age](https://github.com/FiloSottile/age)? I started using it with [agenix](https://github.com/ryantm/agenix), just because [Henrik](https://github.com/hlissner/) started using it in his dotfiles. The [k8s@home template](https://github.com/onedr0p/flux-cluster-template) also eventually [switched from gpg to age](https://github.com/onedr0p/flux-cluster-template/pull/153) so I was already pretty comfortable with using age. TL;DR [Age: the modern alternative to GPG --- nixFAQ](https://nixfaq.org/2021/01/age-the-modern-alternative-to-gpg.html): - Age is presented as a modern alternative to GPG that solves many of its limitations while maintaining security. - Age stands for \"Actually Good Encryption\" and has implementations in Go and Rust for improved security compared to GPG\'s C implementation. - Age uses smaller keys that are easier to store physically and has a simpler interface with no configuration options. - Files can be encrypted for multiple recipients simultaneously using Age. - Age supports encrypting for SSH public keys in addition to its own keys. - Age allows encrypting files for GitHub users by using their SSH keys from their profile. - Age offers a better user experience than GPG while maintaining an equally high level of security. ## Why still org-journal? I\'ve considered using [org-roam-dailies](https://www.orgroam.com/manual.html#org_002droam_002ddailies) instead of [org-journal](https://github.com/bastibe/org-journal). However, after some reflection, I\'m not entirely sure why. [Org Roam supports Age encryption](https://github.com/anticomputer/age.el#org-roam-support-for-age-encrypted-org-files), and [org-journal has several PR fixes](https://github.com/bastibe/org-journal/issues/400) for various issues that have been neglected (this is not a judgment of the maintainer, but problems like [journal files being decrypted whenever the calendar is invoked](https://github.com/bastibe/org-journal/issues/375) are troublesome). I think I was just seeking a quick solution to resume journaling. # Setup Now that you\'ve listened to me ramble on for a bit, here\'s the actual setup. This is using [Doom Emacs](https://github.com/doomemacs/doomemacs). ```elisp title="packages.el" (package! age) ``` ```elisp title="config.el" (use-package! age :init (setq! age-program "rage") :config (setq! age-default-identity "~/.ssh/id_ed25519" age-default-recipient "~/.ssh/id_ed25519.pub") (age-file-enable)) ``` Went with [rage](https://github.com/str4d/rage), because Rust. Also there\'s [pinentry support through rage](https://github.com/anticomputer/age.el#workaround-pinentry-support-through-rage). ```elisp title="config.el" (after! org (setq org-journal-encrypt-journal nil org-journal-file-format "%Y%m%d.org.age") ``` Only thing I had to do then was turn off the built-in gpg support on org-journal, and update the naming scheme to have age as the suffix and it just worked.™ --- # Site v3 Or is it like v5 at this point? Or is it like [v5 at this point](personal-rewrite.org)? In the lead up to [JuliaCon 2023](https://juliacon.org/2023/) I was throwing together the final touches on my presentations and talking to [Teco](https://tecosaur.net/) a lot about his presentation and started digging into the org file he used to create it. I had followed his beamer recommendations in the past, I just had never seen behind the curtain of how he does it. Then it hit me. Why did I never try just using org export for my personal site? I continually searched for ways to write my posts in org-mode, but I always wanted to use the newest and fanciest web framework. Why? I\'m not a web developer. [I know enough to be dangerous](learn-react.org), but the things I was trying to do were vastly over complicated for a personal site. It\'s some static text, and some links. It was never going to need all of the functionality of astro. I mainly wanted to use astro because it would just ship html by default and be fast. Spoiler alert, org-mode does that out of the box. # Scoping what I actually needed 1. I love automated publishing. I literally don\'t think I could run a website or software project without it. The thought of manually walking through the build process manually is like nails on a chalkboard for me. I should just push it to a repo, and boom results. Manual steps are prone to me getting bored halfway through and never finishing it. 2. I hate checking generated files into repos. That means html files, coverted markdown files from org-export. It just clutters the commits and you can\'t follow the history. So from that I planned on: 1. Start with plain html export 2. Get all of the content in the right places 3. Throw some css on there. 4. Then answer the age old question, do you really need JavaScript?[^1] [^1]: I noticed [NvChad](https://nvchad.com/) used [UnoCSS](https://unocss.dev/) and [Solid.js](https://www.solidjs.com/), which seem minimal, while I was looking for a [Doom Emacs](https://github.com/doomemacs/doomemacs) like [Neovim](https://neovim.io/) experience. Planning on using those first. --- # I'm Starting a Writing Streak We'll see how long it lasts this time! This is my second year reading [The Daily Stoic](https://www.goodreads.com/book/show/29093292-the-daily-stoic). Of course with anything that starts on January 1st, I\'m a month behind half way through the year. I was reading May 16th last night and it talked about \"The Chain Method\". Here\'s the explanation: > The comedian Jerry Seinfeld once gave a young comic named Brad Isaac > some advice about how to write and create material. Keep a calendar, > he told him, and each day that you write jokes, put an X. Soon enough, > you get a chain going---and then your job is to simply not break the > chain. Success becomes a matter of momentum. Once you get a little, > it's easier to keep it going. My partner has recently started writing a ton for her new job, and she loves a streak so I suggested we start a writing streak on my [Lego Grad Student](https://brickademics.com/) calendar(Which also happens to still be in May). This is day one. This won\'t always be a blog post, it could be writing in [my slipbox](https://slipbox.edmundmiller.dev/), or reading a paper and taking notes(also in the slipbox). We\'re also counting writing documentation. I\'m hoping the physical calendar and X\'s help! --- # Why I'm hyped about Julia for Bioinformatics Julia might just change the game for bioinformatics. Julia might just change the game for bioinformatics. Bioinformatics is starting to heat up as field, and I think we have a lot of unique issues starting from the days of [perl and the Human Genome Project](https://bioperl.org/articles/How_Perl_saved_human_genome.html). Based off a recent recommendation from my friend [Teco](https://github.com/tecosaur) # Two Language Problem Plenty of people have written about this with Julia. Lots of scientific communities may use a high-level \"scripting\" language to manage their business-logic and then drop into a low-level language such as C/Fortran/Rust to highly optimize the bottle-neck. In bioinformatics we have a 5 language problem. From the perl script written before the first generation of Next Generation Sequencing, to packages like DESeq2 that aren\'t going to leave anyone\'s toolbox anytime soon, but no one is going to rewrite, and instead are going to spend time teaching R to every incoming generation of bioinformaticians just so they can use these influential packages. That\'s not including python, Rust, and every bioinformatician\'s favorite, bash. That\'s not including if you want to contribute to any GATK cli tools written in Java. Now, workflow managers such as [Snakemake](https://snakemake.readthedocs.io/en/stable/) and [Nextflow](https://www.nextflow.io/) have fixed the majority of those. Just stick those scripts in a container that acts as a time-capsule of versions long forgotten and you can use them for just what you need and put them back in closet. But what about new scripts? When you\'re starting a new bioinformatics project in 2022, do you follow in the footsteps of the pioneers that came before you, and learn how to call C functions in an R package? Or would you rather just see this beauty: ## But what about my C functions that are highly optimized? What about our in-house python library that interacts with all of our sample tracking and pulling files temporarily from s3? \"I don\'t have time to figure out how to untangle that mess\" you say. The Julia community wants to keep all of those import scientific scripts written. Packages such as [RCall.jl](https://juliainterop.github.io/RCall.jl/), [PyCall.jl](https://www.juliapackages.com/p/pycall), and plenty of others housed under [Julia Interop](https://github.com/JuliaInterop) allow for legacy code to fit right into new code. As scientists we are constantly standing on the shoulders of giants, and Julia is enabling that. # [BioJulia](https://biojulia.net/) Julia has a wonderful open-source structure. Since it was created only a decade ago there\'s not a lot of legacy things to maintain and everything can be fresh and inclusive. Take for example the heavy use of organizations to house these code repositories, which prevents packages from going unmaintained and orphaned when the creator moves on. While it\'s a relatively small community they\'ve covered a wide range of the various file types that we deal with on a daily basis. I hope to dog food most of the packages to fill in some missing pieces, and improve documentation and create some content for the community. The part that I found difficult was finding the people based on the website. I joined the gitter to ask a question, but luckily someone saw the message and told me that most of the real-time chat happens in [the #biology channel in the Julia slack](https://julialang.slack.com/archives/CAKKFNYLD). # Designed for Scientific Computing It seems like bioinformaticians always want to do things in the most efficient path, but solving for a different variable than most. [In a recent episode of screaming in the cloud](https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/quantum-leaps-in-bioinformatics-with-lynn-langit/), [Lynn Langit](https://lynnlangit.com/) mentioned that in finance, they care about getting the results as quickly as possible. The cost of the computing isn\'t a factor. In bioinformatics, they\'re trying to solve for both time and cost, trying to find the local minimum between both. There are plenty of time results don\'t need to instant and waiting a few days to stretch a grant out is a necessary evil. Julia is fast, and that\'s been said numerous times, so you\'re probably guessing I\'m going to say you can save money by decreasing your compute time. While that\'s true the piece that I think people aren\'t solving for in that equation is developer time. Julia was designed from the ground up as a general programming language for scientists ([Why We Created Julia](https://julialang.org/blog/2012/02/why-we-created-julia/)). > We want a language that\'s open source, with a liberal license. We > want the speed of C with the dynamism of Ruby. We want a language > that\'s homoiconic, with true macros like Lisp, but with obvious, > familiar mathematical notation like Matlab. We want something as > usable for general programming as Python, as easy for statistics as R, > as natural for string processing as Perl, as powerful for linear > algebra as Matlab, as good at gluing programs together as the shell. > Something that is dirt simple to learn, yet keeps the most serious > hackers happy. We want it interactive and we want it compiled. That sounds like a bioinformatician\'s dream to me! Think of all of the time we can save on developer experience, not trying to hack out some extra speed, or fixing broken dependencies (or a complete lack of specified dependencies!). Allowing legacy code to be treated like the crown jewel in our metaphorical software crown, and wrapping it in some gold Julia code like it deserves. # Call to action I plan on working to increase the visibility into Julia specifically for bioinformatics. I always love this Venn diagram to explain to people what bioinformatics is. I think it\'ll be natural to cover [JuliaData](https://github.com/JuliaData/), [JuliaStats](https://juliastats.org/), and [BioJulia](https://biojulia.net/) to cover all three and show people how the three intersect. ## How to get started with Julia - [Attend JuliaCon 2022](https://juliacon.org/2022/tickets/) (It\'s online and free to attend!) - [Listen to the Talk Julia Podcast](https://www.talkjulia.com/) - [Go through some courses on JuliaAcademy](https://juliaacademy.com/) - Check out the work by [Logan Kilpatrick, the Julia Dev Community Advocate](https://www.logankilpatrick.com/) ### 2023 Update - [GitHub - BioJulia/BioTutorials: Tutorial Notebooks of BioJulia](https://github.com/BioJulia/BioTutorials) - New Documenter.jl Docs! --- # Thoughts on NixOS Highlighting its customization, stability, and how it simplifies development environments. NixOS is a dream. It allows you to program your os by declaring what you want in it. And then you can declare how you want your packages built. Most of the time you just want the default, but for example ```nix (polybar.override { mpdSupport = true; pulseSupport = true; nlSupport = true; }) ``` It gives you the power to customize things when you want. And then contributing packages is more developer focused, it\'s just a PR away. It makes it really simple to mix bleeding edge, with stable. The rollbacks are another big initial selling point that you kind of forget about because things end up being so stable but they\'re the best. Basically in your grub you can select any generation you\'d like, so in case you wanted to try out a new kernel and that break you just rollback to the previous generation. The drawbacks are you need to understand a bit of functional programming, and then on top of that you have to learn nix which is a dsl. There\'s nothing wrong with the language, just that it\'s another thing you have to learn. It may also kill your hobby of ricing. It makes it so quick to reproduce a setup, that you can spend time actually thinking about the artistic portion and not symlinking config files, and getting the right package version on ubuntu. I follow a guy\'s dotfiles and he created a modular system, so your theme separate from the logic that sets up your WM, so you can carry your rice across WM easily and themeses across a wm. [Link to my Dotfiles](https://github.com/Emiller88/dotfiles) Guix, cuts out the dsl issue with nix as the language. I love a good lisp personally, and it\'s definately easier to pick up than nix. Nix/OS is also a pretty old project, so it\'s grown over time, so the tooling was build up over time so there\'s a bunch of different tools like nix-shell ,nixos-install, nixos-rebuild, nix-env that are all at the first level where as guix has all of those things under guix so it\'s better for new people to discover commands from. Again the main issue is they\'re going to die on the FOSS hill, but I think that\'ll get fixed with private package repos. Now that\'s just all for the OS. They can both be used for dev-tools which is where they really shine imo. You can use nix/guix on MacOS and any distro of your choosing. You can create your builds and dev environment for any language in them. So for example, you\'re just trying out python for the first time. You\'re overwhelmed with the 20 different ways to set up a developement environment. With nix it\'s just ```nix let pkgs = import {}; nanomsg-py = .... build expression for this python library; in pkgs.stdenv.mkShell { buildInputs = [ pkgs.pythonPackages.pip nanomsg-py ]; shellHook = '' alias pip="PIP_PREFIX='$(pwd)/_build/pip_packages' \pip" export PYTHONPATH="$(pwd)/_build/pip_packages/lib/python2.7/site-packages:$PYTHONPATH" unset SOURCE_DATE_EPOCH ''; } ``` Another example that I had recently, was I needed an older version of node, but I didn\'t want to clutter up my path with multiple node packages I just wanted to build a project really quick with `node_10`{.verbatim}. All it was is `nix-shell -p node_10` and boom I had a shell with `node_10`{.verbatim} installed. So then you can combo the power of nix-shell with direnv, to automagically switch environments based on what project your in, and it\'s easy for your team to replicate also, and let\'s you be more language-agnostic because while tooling is awesome when it\'s good(rust for example, or node possibly) when it\'s scary or the community get\'s fragmented(python) it\'s a problem. If anyone made it this far and you want more come hang out [in the Doom Emacs Discord](https://doomemacs.org/discord), you don\'t have to be an Emacs user even, but you might end up one. --- # A New Vue on Life I've been wanting to redo my personal site since I tried to add my I\'ve been wanting to redo my personal site since I tried to add my presentations to my hugo based site and struggled. Something that looked cool. Something that was interesting. At first I started looking at static site generators and found a great site, [staticgen](https://www.staticgen.com/), that lists all the possible static site generators. I thought I wanted one written in a cool language, like a lisp or Haskell, so I went down all the lisps first, and I realized my first real requirement was simple CI deployment. I like GitLab CI for it\'s control, [zeit now](https://zeit.co/) as a close second for it\'s ease, and then maybe Netlify. The lisp family didn\'t have anything quick to deploy and their installation stories were equally poor. So I jumped to Hakyll, there\'s an [example](https://gitlab.com/pages/hakyll) for GitLab Pages so I got started on that. And I thought I wanted to build it with Nix and found a few brave souls who had ventured down that path. But I figured I\'ll just try the example first, why not. The pipeline took 30 minutes to build, because the Haskell packages had to be installed, which would be fine, if the cache lasted longer than a week. I then realized that I didn\'t really want just a static site generator, they are using these fantastic powerful programming languages to just generate static pages, just a glorified Makefile. And I didn\'t want to use someone else\'s theme, but I didn\'t want to write my own in the generators specific formatting. It felt limiting, but I didn\'t want to reinvent the wheel. I didn\'t like making websites. Maybe I\'m just not good at design, but I\'d rather focus on the content than trying to center a div. So I thought maybe it was just javascript. So I tried Elm. I settled on [elm-pages](https://elm-pages.com/) , what finally sold me over elm-static was the [introduction blog post](https://elm-pages.com/blog/introducing-elm-pages/) and it insightful into the JAMstack (JavaScript, APIs, and Markup). I had used [gatsby](https://www.gatsbyjs.org/) to rewrite the [UTDallas Rugby Page](https://www.utdallasrugby.org/) with [Tristen Even](https://www.tristeneven.com/) so the pieces were starting to fall together in my head about what JAMstack is about. So I got it up and running. And it sat there because I don\'t elm. It looks fantastic checks all the functional buzz word bingo card spots. But my list of things to learn is taking on more water than I can bail out, and elm kept sinking to the bottom. Fast forward a couple of weeks, to [ETHDenver](https://www.ethdenver.com/) and I met Jeremy, a friend of a friend and one of the creators of [loft.radio](https://loft.radio/), a lofi hip hop radio that you can tip the artists in ETH. We we\'re chatting tech, as one does at a hackathon, and he was raving about Vuejs, which he used to build loft. I had heard about it from another friend and decided to checkout after I finished my weekend with react. I had the perfect excuse to dip my toes in when I needed to build some documentation for work, and settled on [VuePress](https://vuepress.vuejs.org/). I think all documentation should be written in markdown, and then the site generated to prevent vendor lock-in. It\'s been a great experience so far and has a ton of features out of the box. But like Vue, it let\'s you walk into the water, rather than jump into the deep end like React. Just my personal experience. That small brush inspired me to dust off an old side project, [Multiverse](https://multiverse.gg/). It\'s my buzzword labor of love, Nextjs, Tailwindcss, Typescript, Ava, Graphql, and Stripe. I had patched into together like Frankenstein from several [examples](https://github.com/zeit/next.js/tree/master/examples), but I got hung up between Typescript and Ava and getting it to build and test those. In a perfect storm the Vuejs documentary came out at the same time so I rewrote my tiny progress in [Nuxtjs](https://nuxtjs.org/). I think this quote from the Vuejs docs section comparing it to other frameworks is my experience exactly > For many developers who have been working with HTML, templates feel > more natural to read and write. The preference itself can be somewhat > subjective, but if it makes the developer more productive then the > benefit is objective. I\'ll just cover my experience really quick. Vue definitely takes the developer experience as a priority and that reflects on the frameworks created on top of it. Just give `vue create hello-world`{.verbatim} or `npx create-nuxt-app `{.verbatim} a shot and you\'ll see the difference from the blank canvas that `creat-react-app`{.verbatim} gives you. From `create-nuxt-app`{.verbatim} I just had to add the [apollo-module](https://github.com/nuxt-community/apollo-module) and I had most of my wish list taken care. In the past I had used [prisma](https://www.prisma.io/) for a quick DB with Graphql on top, but I found they had updated their docs to offer more, but I wanted something simpler. Enter, [Hasura](https://hasura.io/), one click deploy to Heroku. A docker image for when you want to move somewhere else. Makes thinking about the back end an after thought. I made more progress in the past two weeks than in the past 9 months. I\'m just missing Typescript and Stripe so far, and I\'m assuming Stripe will be a breeze. Maybe I\'m just not a clever enough developer for React. I love functional programming and React Hooks are just sexy and they still nearly lure me back in. But Vue gave me a new light on web development and programming languages. Vue gave me a paintbrush to finally appreciate design and the joy of making a web page come to life rather than wrangling JavaScript and never getting anywhere. I realized what the advice a language is a tool to get a job done meant. The paint from Vue may crack after a while, as apposed to the copper statue that Elm and similar languages promise. While it\'s great, it won\'t matter if I never make the statue because I can\'t learn metallurgy at a craft store. # Tailwind Refactor I also wanted to take this chance to explore [tailwindcss](https://tailwindcss.com/). I had exactly 0 experience with CSS before this and fought with any React material design component I tried even get to agree with flexbox. I had seen tailwind from [Zamansky\'s ClojureScript tutorial](https://www.youtube.com/watch?v=_CTTbC6owS0), and [Aria\'s](https://github.com/ar1a) Twitter clone [Catter](https://catter.netlify.com/). So I took rewriting the default Gridsome as the perfect way to get accustomed to using it, without getting caught up in the design aspects because I haven\'t gotten to read [Refactoring UI](https://refactoringui.com/) yet. As of [3e57e24661](https://github.com/Emiller88/edmundmiller.dev/tree/3e57e2466116fc260c077239d5cfdf4c0063ee40) I completed the rewrite. To speak on my experience with tailwind, it really made it easy to focus on what I wanted to do, rather than how to do it. I had a friend show me webflow recently, and I think tailwind is even simpler once you get familiar with the utility-classes it provides. The original was a beautiful intricate work of CSS, that I honestly had no interest in figuring out. The one issue I did run into was the way Gridsome just returned a `v-html="$page.post.content"` so I found the way [Gridsome Portfolio Starter](https://gridsome.org/starters/gridsome-portfolio-starter/) handled it\'s markdown and then threw some tailwind in to deal with the theme-switcher specific stuff and ended up with [github-markdown.css](https://github.com/Emiller88/edmundmiller.dev/blob/3e57e2466116fc260c077239d5cfdf4c0063ee40/src/assets/css/github-markdown.css) ## [TODO]{.todo .TODO} Theme Switcher {#theme-switcher} # [TODO]{.todo .TODO} Semantic Versioning {#semantic-versioning} # [TODO]{.todo .TODO} Presentations {#presentations} --- # Learn Just Enough React to Move Out of your Parent\'s House A Gen-Z's Guide to the 2019 Job Market. # Introduction I recently was helping a friend start building a website, and in hopes of boosting his resume, chose to build it in [React](https://reactjs.org), and because I think the best way to master something is to teach it. I had picked up the frame work through work, and I recommend [React Holiday](https://react.holiday) for anyone trying to pick it up with just a few minutes a day to spare. The title came from the number of friends that I have that have learned the popular framework and then been able to move out of their parent\'s house. I also was inspired by the [Hacker News Hiring Trends](https://www.hntrends.com) which claims react has been the most popular skill requested for 22 months and made up 28% of all job postings in March 2019. Because of the surplus of tutorials that seem to teach you `Just Enough`{.verbatim} ™. Hopefully this post will be enough to get you going from zero to out of your parent\'s house. ## Install git You\'ll want install `git`{.verbatim} next. Follow the instructions at this [site to install](https://git-scm.com) for your platform. ## VS Code I personally don\'t use VS Code, I use [Doom Emacs](https://github.com/hlissner/doom-emacs) but that\'s a whole other beast. With that disclaimer I recommend VS Code to friends new to programming because it is the popular GUI editor of the hour. You can download [VS Code](https://code.visualstudio.com). ## Fira Code A fun little thing to make you feel like a special snowflake is to change your font in your editor. I like to recommend `FiraCode`{.verbatim} because it has ligature support with overwhelming you with options. You can see how to install it on your system and add it VS Code at the [FiraCode installation guide](https://github.com/tonsky/FiraCode/wiki). ## Install nodejs _Node.js is a JavaScript runtime built on Chrome\'s V8 JavaScript engine_. is the description their site gives. We\'ll cover more on JS later. The link to the [Download page](https://nodejs.org/en/download/). ## Installing Extensions One of the perks of VS Code is that it makes install extensions simple. They can be installed by clicking the square icon on the left side of VS Code. The list of what I recommend for following this guide: - [vscode-icons](https://marketplace.visualstudio.com/items?itemName=vscode-icons-team.vscode-icons) - [Bracket Pair Colorizer](https://marketplace.visualstudio.com/items?itemName=CoenraadS.bracket-pair-colorizer) - [Debugger for Chrome](https://marketplace.visualstudio.com/items?itemName=msjsdiag.debugger-for-chrome) - [GitLens](https://marketplace.visualstudio.com/items?itemName=eamodio.gitlens) - [One Dark Pro](https://marketplace.visualstudio.com/items?itemName=zhuangtongfa.Material-theme) - [Settings Sync](https://marketplace.visualstudio.com/items?itemName=Shan.code-settings-sync) - [npm](https://marketplace.visualstudio.com/items?itemName=eg2.vscode-npm-script) - [React Snippets](https://marketplace.visualstudio.com/items?itemName=dsznajder.es7-react-js-snippets) - [VS Intellicode](https://marketplace.visualstudio.com/items?itemName=VisualStudioExptTeam.vscodeintellicode) - [IntelliSense for CSS](https://marketplace.visualstudio.com/items?itemName=Zignd.html-css-class-completion) - [Path Intellisense](https://marketplace.visualstudio.com/items?itemName=christian-kohler.path-intellisense) You should glance over each and see what they give you. ### [prettier](https://marketplace.visualstudio.com/items?itemName=esbenp.prettier-vscode) `prettier`{.verbatim} formats the code you on save. Follow the instructions on the page to set it up. ### [EditorConfig](https://marketplace.visualstudio.com/items?itemName=EditorConfig.EditorConfig) Another way to keep your code clean and set a standard. Can you tell I like clean code? ### [eslint](https://marketplace.visualstudio.com/items?itemName=dbaeumer.vscode-eslint) `ESLint`{.verbatim} tells you when you\'ve done something wrong and sometimes can fix it for you. ## Integrated Terminal If you know how to get a terminal on your system already awesome. If you don\'t know what I\'m writing about, we\'re just going to use the integrated terminal. [A quick doc on it](https://code.visualstudio.com/docs/editor/integrated-terminal), I recommend you use `git bash`{.verbatim} if you\'re on [windows](https://code.visualstudio.com/docs/editor/integrated-terminal#_windows). # GitHub ## Create an Account If you don\'t have one yet that\'s fine. It should be self explanatory. Bonus if you\'re a student, sign up for the [Student Pack](https://education.github.com/pack). I am not spelling out a majority of this tutorial because of the various different platforms people are one would make it quite long, and because part of developing software is learning to read the friendly manual. Think of it more of a checklist. ## Create a Repo I would be doing the docs at GitHub a disservice if I tried to take their tutorial so here is [Creating a new repository](https://help.github.com/en/articles/creating-a-new-repository). We won\'t be covering a personal site today so name it whatever you like. I\'ll be referring to it as `example-site`{.verbatim}. ## Clone the repo Using your `Integrated Terminal`{.verbatim} follow GitHub's [repository cloning guide](https://help.github.com/en/articles/cloning-a-repository) to clone your site. In my case it would be: ```bash git clone https://github.com/edmundmiller/example-site ``` # Getting the Site set up Now that I\'ve bored you with all of the tooling, or if you enjoyed it, we\'re on to the real work. ## Create React App Is a great utility to get your up and running with `React`{.verbatim} ```bash npx create-react-app my-app cd my-app npm start ``` If you installed `nodejs`{.verbatim} correctly earlier this should go off without a hitch and you should have a browser popup with your site. This is a `local`{.verbatim} site that hot reloads whenever you edit anything in the project so you can get feedback if your change is correct quickly. ## GitHub Pages Follow the [Procedure](https://github.com/gitname/react-gh-pages#procedure), you should be able to skip to step 3. replace `react-gh-pages`{.verbatim} with `example-site`{.verbatim} or whatever you chose. ## CircleCI Lastly, we\'ll setup a CI/CD pipeline to automatically deploy and build your site whenever you push code to master. You\'ll want to [create an account](https://circleci.com) and link your GitHub. We\'ll be following this [blog post](https://circleci.com/blog/automate-your-static-site-deployment-with-circleci/). Here is the `.circleci/config.yml`{.verbatim} you\'ll need to add to your project. ```yaml version: 2 jobs: build: docker: # specify the version you desire here - image: circleci/node:lts # Specify service dependencies here if necessary # CircleCI maintains a library of pre-built images # documented at https://circleci.com/docs/2.0/circleci-images/ # - image: circleci/mongo:3.4.4 working_directory: ~/repo steps: - checkout # Download and cache dependencies - restore_cache: keys: - v1-dependencies-{{ checksum "package.json" }} # fallback to using the latest cache if no exact match is found - v1-dependencies- - run: npm install - save_cache: paths: - node_modules key: v1-dependencies-{{ checksum "package.json" }} # run tests! - run: npm run test - deploy: name: deploy to GH-Pages command: npm run deploy ``` # React It\'s about time we actually talked about `React`{.verbatim}. As you can see though a good chunk of development is just setting up the project. There\'s obviously the link to the official documentation that comes in the `create-react`{.verbatim} starter page which I recommend you read. But now that we\'re to the actual meat I\'ll take you through a few things. ## React bootstrap If you\'ve ever seen a basic website recently it might be made with bootstrap. It was recreated for use with [React](https://react-bootstrap.github.io/getting-started/introduction). [To get started with it](https://react-bootstrap.github.io/getting-started/introduction) run the following and then follow the docs. ```bash npm install react-bootstrap bootstrap ``` --- # Emacs Email Config Setting up Email in Emacs with mbsync, and Notmuch. # Introduction Like most people these days I have quite a few email addresses, personal, old personal email, and school. In the past I've tried using Emacs to manage all of these in one place. However, when I added a work email account into the mix that didn't have `IMAP` enabled it was finally enough to make me go back to using the web clients. Recently I started to get the itch to bring the config into my [dotfiles](https://github.com/Emiller88/dotfiles/tree/master/shell/notmuch), since the biggest pain was setting it all up on a new computer and making it feel fragile. ## Guides: - [Notmuch of a mail setup Part 1- mbsync, msmtp and systemd](https://bostonenginerd.com/posts/notmuch-of-a-mail-setup-part-1-mbsync-msmtp-and-systemd/) - - ## Notmuch config of my dotfiles: - # Fetching the Mail In the past I\'ve tried [offlineimap](https://github.com/OfflineIMAP/offlineimap) but it\'s slow when you have to pull down years and years of email. So I first started with [mbsync](https://wiki.archlinux.org/index.php/Isync)(I\'ve used the arch wiki because it explains more than the actual documentation in my opinion). To get that set up: ```bash sudo apt install isync ``` and then add a `~/.mbsyncrc` file and configure it according to one of the above guides or the [arch wiki](https://wiki.archlinux.org/index.php/Isync). I will try to avoid going over things that I just copied and pasted and they worked, in order to keep this short. An issue that I ran into later down the road was that [mbsync](https://wiki.archlinux.org/index.php/Isync) doesn\'t sync just any flag/label/tag back to Gmail. Enter [gmailieer](https://github.com/gauteh/gmailieer). Which was more of a pain to setup. To install from [Ubuntu's gmailieer package](https://launchpad.net/ubuntu/+source/gmailieer) ```bash sudo apt-get install gmailieer ``` The chicken and the egg problem with `gmailieer` is that it needs `notmuch` set up in the `.mail` dir. So when I was setting `gmailieer` up, `notmuch` was already up and running for me. We\'ll come back to it. With `gmailieer` the important part, is if I check my email on my phone, another computer, or Emacs, it will update as read, archived, deleted, and most importantly the tags will be the same on it. For an added bonus it works for my work account that uses Gsuite since it uses their API and not IMAP. Now one key issue I had with `gmailieer` and my setup is that it depends on the `python3-notmuch` package, which is in the Ubuntu repos but not in `pypi`. So this makes it necessary to use a different version of python than Conda 3.6(default in 18.04 but you can use the `python-notmuch` package if you\'re on 16.04) I had to do the following(I use [pyenv](https://github.com/pyenv/pyenv) to manage different python versions). ```bash sudo apt install python3-notmuch cd ~/.mail pyenv local 3.7.0 # or whatever version you like ``` We need to initialize the notmuch database. ```bash sudo apt-get install notmuch cd ~/.mail notmuch ``` Here you\'ll be prompted with some questions to get your config going. Below is an example of what it\'ll look like afterwards ```toml [database] path=/home/$USER/.mail [user] name=Edmund Miller primary_email=USER1@gmail.com other_email=USER3@gmail.com;USER2@school.edu; [new] tags=new ignore=Trash;*.json; [search] exclude_tags=trash;deleted;spam; [maildir] synchronize_flags=true ``` # Checking Mail in the Background Next we need to setup the mail to be ran in the background every so often. This [BostonEnginerd's systemd mail setup guide](https://bostonenginerd.com/posts/notmuch-of-a-mail-setup-part-1-mbsync-msmtp-and-systemd/%0A) has a great setup for it. ```bash mklink checkmail.service $XDG_CONFIG_HOME/systemd/user/checkmail.service mklink checkmail.timer $XDG_CONFIG_HOME/systemd/user/checkmail.timer ``` And start and enable the timer ```bash systemctl --user start checkmail.timer systemctl --user enable checkmail.timer ``` The `checkmail.service` calls `checkmail.sh` ```bash #!/usr/bin/env bash STATE=$(nmcli networking connectivity) if [ $STATE = 'full' ]; then echo "Syncing email1" cd /home/emiller/.mail/email1/ gmi sync echo "Syncing email2" cd /home/emiller/.mail/email2/ gmi sync echo "Syncing work" cd /home/emiller/.mail/work/ gmi sync echo "Checking school" # Non gmail email mbsync -V school exit 0 fi echo "No internet connection." exit 0 ``` The `gmi sync` command does a `push` followed by a `pull` so the tags from the local overwrite anything that\'s on the remote. So later we\'ll write rules to tag the new mail coming in. # Tagging the Mail The next step is to tag the mail. For that I use `notmuch` I tried `mu` in the past but it works by moving the emails into various dirs instead of just tagging them and I found it messed with how the remote emails were treated too often. Gmailieer pulls the tags down by default. But if we want to tag our mail locally we\'ll need to expand `checkmail.sh`. [afew](https://github.com/afewmail/afew) is another option for more elaborate initial tagging, but I didn\'t want to have more dependencies. ```bash #!/usr/bin/env bash STATE=$(nmcli networking connectivity) function tagMail { echo "Running tag additions to tag new mail" # github notmuch tag +github -- from:notifications@github.com AND tag:new notmuch tag +github -- from:noreply@github.com AND tag:new notmuch tag -inbox -- tag:github AND tag:new # CI notmuch tag +CI -- from:builds@travis-ci.com AND tag:new notmuch tag +CI -- from:builds@circleci.com AND tag:new notmuch tag -inbox -- tag:CI AND subject:Passed # Mailing Lists notmuch tag +list/emacs -inbox -- from:help-gnu-emacs-request@gnu.org AND tag:new notmuch tag +list/haskell -inbox -- from:info@haskellweekly.news AND tag:new notmuch tag +list/IPFS -inbox -- from:newsletter@ipfs.io AND tag:new notmuch tag +list/rust -inbox -- from:twir@rust-lang.org AND tag:new notmuch tag +list/nixos -inbox -- from:domen@enlambda.com AND tag:new # Remove new notmuch tag -new -- tag:new } if [ $STATE = 'full' ]; then echo "Syncing email1" cd /home/emiller/.mail/email1/ gmi sync echo "Syncing email2" cd /home/emiller/.mail/email2/ gmi sync echo "Syncing work" cd /home/emiller/.mail/work/ gmi sync echo "Checking school" # Non gmail email mbsync -V school echo "Running notmuch new" notmuch new echo "Tagging mail" tagMail exit 0 fi echo "No internet connection." exit 0 ``` So what we\'re doing here is first calling `notmuch new` which tags everything according to this section of the config. Which just tags everything with `new` and ignores anything with the `Trash` tag. ```toml [new] tags=new ignore=Trash;*.json; ``` # Deleting Email Notmuch by default doesn\'t tag things with `+trash` which makes gmail move the emails to the trash. Here\'s a snippet that does that. I have this bound to `d`. ```elisp (defun +notmuch/search-delete () (interactive) (notmuch-search-add-tag (list "+trash" "-inbox" "-unread")) (notmuch-tree-next-message)) ``` # WIP: Sending Email > WIP: Currently I can only get this to work with my primary email > address. To set this up we\'ll need to get started with pass. I suggest you have a look at the [Doom Notmuch module](https://github.com/hlissner/doom-emacs/blob/develop/modules/email/notmuch/config.el) if you\'re not using Doom to give you an idea of any features you need to setup. First setup `~/.msmtprc` ```toml # Set default values for all following accounts. defaults auth on tls on tls_trust_file /etc/ssl/certs/ca-certificates.crt logfile ~/.msmtp.log # Gmail account gmail host smtp.gmail.com port 587 from USER1@gmail.com user USER1 passwordeval pass mail/USER1 ``` Then we\'ll setup `pass`. ```bash pass init pass insert mail/USER1 ``` And type in the password. You should be good to go and when in `notmuch` hit `C` and the `C-c C-c` to send and `C-c C-k` to cancel. ``` C Compose new mail. ``` --- # CTRL to Caps Lock How to change your Capslock key to Ctrl. With a bonus ESC on tap. I grew up in a household where `CTRL` has been mapped to `Caps Lock` on every computer. It was just something I had never questioned until recently. As a heavy `CTRL` user between Emacs and terminal apps now, I understand.I thought this would be fitting for a first blog post as kind of an origin story. So I got my first mechanical keyboard about a year and a half ago, and since it 's powered by [qmk](https://docs.qmk.fm/#/), I decided to replace where caps lock might have been on my planck, to `ESC` to make Vim and evil a little easier instead of using `jk` or other ways to switch back to normal mode. Fast-forward to recently and my mechanical keyboard family has grown with the addition of an ergodox-ez. The default bindings had a lot of keys that with a quick tap was a letter, and a hold became a modifier key. So I thought it might be useful to do this with `ESC` and `CTRL`. This is the line to make it happen with [qmk](https://docs.qmk.fm/#/). ```c CTL_T(KC_ESC) ``` But it wasn\'t just enough to have it on my keyboards. I needed this magic everywhere. After mentioning what I had done on the Doom Emacs discord one of them found a [guide by Danny Guo](https://www.dannyguo.com/blog/remap-caps-lock-to-escape-and-control/) to set this up on most systems. [Instructions for Ubuntu](https://www.dannyguo.com/blog/remap-caps-lock-to-escape-and-control/#xcape) however, under `gnome tweaks` under keyboard, additional layout options, Ctrl position, it was `Caps Lock as Ctrl`. I now have ESC/CTRL set to caps lock on keyboards not powered by qmk, such as my laptops built in keyboard.