# Sharif Haason — full writing corpus This is the complete public writing archive in Markdown. The canonical page for each article and its structured JSON representation are included in the metadata block. ## A New Chapter: Graduation and a New Role at Monark - URL: https://sharifhsn.dev/blog/linkedin-2026-04-01-a-new-chapter-graduation-and-a-new-role-at-monark/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2026-04-01-a-new-chapter-graduation-and-a-new-role-at-monark.json - Description: I'm thrilled to make two major announcements: I have completed my Master's in Financial Engineering from Stevens Institute of Technology, and I am starting a new role at Monark Mar… - Date: 2026-04-01 - Exact published timestamp: 2026-04-01T14:42:17.714Z - Topics: Career, Personal - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7445118347752402944/ I'm thrilled to make two major announcements: I have completed my Master's in Financial Engineering from Stevens Institute of Technology, and I am starting a new role at Monark Markets as a Product Engineer! I'd like to specially thank Dr. Ionut Florescu, Dragos Bozdog, Zheng Xing, and Zhenyu Cui for their mentorship at Stevens, as well as my coworkers at the Hanlon Financial Systems Center at Stevens Institute of Technology, my professors, my classmates in my program, and everyone else at Stevens who has helped me on my journey. Additionally, I never could have done this without the support of my friends and family, who I cherish always. I am excited to take on new challenges at Monark Markets and I believe their product is a complete hit in the making. If anyone is interested in investing in private markets like SpaceX or Anthropic through retail brokerages, please look into what we're doing. ## The December Fed Meeting: What to Watch - URL: https://sharifhsn.dev/blog/linkedin-2025-12-10-the-december-fed-meeting-what-to-watch/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-12-10-the-december-fed-meeting-what-to-watch.json - Description: The December Fed meeting is today! Despite the worrying lack of data, the market is confident that rates will be cut by 25 bps, with the FedWatch tool giving a 87.6% chance for a c… - Date: 2025-12-10 - Exact published timestamp: 2025-12-10T11:00:16.071Z - Topics: Central Banking, Monetary Policy - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7404475031784730625/ The December Fed meeting is today! Despite the worrying lack of data, the market is confident that rates will be cut by 25 bps, with the FedWatch tool giving a 87.6% chance for a cut. With PCE inflation at expectations and no surprises from JOLTS, all indicators are pointing to a cut. But there's more that meets than eye to this situation. This is Jerome Powell's last meeting as Fed Chair, with his successor already decided, though not announced. Wall Street is speculating that Kevin Hassett, current director of the National Economic Council, will replace Powell. He is qualified, which is more than can be said for some of Trump's other appointments, but he is also a MAGA loyalist, which may present issues for the Fed's neutrality going forward. While a cut is expected, close attention should be paid towards Powell's rhetoric during his remarks and responses to Q&A after the meeting, as they will signal what the rest of the FOMC is thinking about the economy. This could be the last meeting before those carefully chosen words are replaced by partisan rhetoric. I'll close this post out with a commendation of Jerome Powell. He fearlessly commanded the Fed through the COVID crisis, the subsequent inflation, and immense pressure from the White House, never wavering in his resolve. The Fed is worse off without him at the helm. What do you think? Will the Fed cut rates? ## Why 50-Year Mortgages Fall Short - URL: https://sharifhsn.dev/blog/linkedin-2025-11-15-why-50-year-mortgages-fall-short/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-11-15-why-50-year-mortgages-fall-short.json - Description: Federal Housing Finance Agency director Bill Pulte recently floated a big potential change to the mortgage market: 50 year mortgages. Unfortunately, these aren’t realistic. - Date: 2025-11-15 - Exact published timestamp: 2025-11-15T14:30:01.051Z - Topics: Housing, Public Policy - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7395468120376627200/ Federal Housing Finance Agency director Bill Pulte recently floated a big potential change to the mortgage market: 50 year mortgages. Unfortunately, these aren’t realistic. Typical mortgages last 30 years. Extending them to 50 years would have the benefit of lower monthly payments for borrowers. However, the idea has several issues. The basic idea is a non-starter from the mortgage holder point of view. They need to hedge away their interest rate risk, or “duration”. If rates rise, then the value of the loan goes down. Therefore, the mortgage holder will short Treasury bonds of the same maturity to make themselves duration neutral. If interest rates move up, then both the loan and the treasury bond will lose value equally if balanced appropriately. The mortgage holder will only profit off of the spread between the Treasury yield and their mortgage rate, and be protected from moves in the interest rate.* But a 50 year Treasury bond doesn’t exist, and it’s doubtful that one would ever be issued. There is enough demand for lower terms that a 50 year Treasury bond would carry an illiquidity premium, on top of the extreme inflation premium. Unless the national deficit balloons massively, there won’t be a funding need for it. Even if this bond did exist, a 50 year mortgage would be horrible for borrowers, causing them to pay significantly more interest over the period of the mortgage. See the image below for a comparison of a $400k principal mortgage. The 50 year rate is guaranteed to be higher than the 30 year rate to compensate for the increased duration, credit, and inflation risk. Paying an extra $500k in interest is extreme, and dangerous when most people don't really understand interest to begin with. Pulte has pulled back from his enthusiasm on the 50 year mortgage after the market responded negatively. However, given the explicit cosign of President Trump, we might continue to hear about this… or they’ll move onto a new product, like portable mortgages. But that’s a subject for another day. Read more at Bloomberg: [https://lnkd.in/edU-T2Zx](https://lnkd.in/edU-T2Zx) Trump’s 50-Year Mortgage Loses Steam as Industry Questions Costs *Treasury bonds do not make a perfect hedge, because mortgages can be refinanced, introducing negative convexity (short gamma) into the asset. Portfolio managers will long gamma (e.g. a payer swaption) to neutralize this cost. The same liquidity problems exist. ## Jefferies and First Brands: What the Trade Debt Means - URL: https://sharifhsn.dev/blog/linkedin-2025-10-09-jefferies-and-first-brands-what-the-trade-debt-means/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-10-09-jefferies-and-first-brands-what-the-trade-debt-means.json - Description: "Jefferies Fund Has $715 Million in First Brands Group, LLC’ Trade Debt"... but what does that exactly mean? - Date: 2025-10-09 - Exact published timestamp: 2025-10-09T14:28:33.464Z - Topics: Credit, Financial Markets - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7382059401983197184/ "Jefferies Fund Has $715 Million in First Brands Group, LLC’ Trade Debt"... but what does that exactly mean? This Bloomberg headline is causing market turmoil about Jefferies, but let's cut through the noise and disambiguate what kind of exposure Jefferies actually has here. [1] First, a quick summary of the story so far. First Brands is an auto parts conglomerate which expanded aggressively by taking on debt through trade financing, which allows the company to be immediately compensated for future receivables by a fund (in this case, Point Bonita Capital) which takes on that risk. When Jefferies pitched a $6 billion refinancing deal to First Brands in July, it triggered scrutiny into First Brands' books, causing its slow implosion into bankruptcy. (The full story of First Brands' bankruptcy is still unfolding and is a very interesting read [2]) Jefferies was quick to open their books and clear up their exposure to First Brands. In the process, they showed the kind of decisions that led to this bankruptcy. In one instance, Point Bonita accepted "side letter" deals with First Brands which allowed them to raise funds at a higher interest rate beyond what the loans allowed, increasing financing costs for Point Bonita. [3] This event shows the risk of trade financing, especially in the dark world of private credit. I've created a diagram showing how Jefferies is exposed. In short, they have a $113MM equity stake in Point Bonita Capital, which is the fund exposed to the eye-popping $715MM in receivables. But according to analysts at Morgan Stanley, estimated losses to Jefferies top out at about $45MM, which is definitely manageable. Some investors, like BlackRock, have pulled out of Point Bonita as a result, causing Jefferies stock to tumble by ~8% yesterday [4]. But I think this fear is unwarranted. Jefferies is in the market for high risk, high reward borrowers, and blowups like this are a cost of doing business. Headline hysteria might cause their stock price to fluctuate, but Jefferies' meteoric rise from boutique to multinational bank is no accident. When you play with fire, you're bound to get singed now and then... but I say "let them cook!" What do you think? How big of a deal is the First Brands' bankruptcy for Jefferies? [1] [https://lnkd.in/eD89yvGd](https://lnkd.in/eD89yvGd) Jefferies Fund Has $715 Million in First Brands' Trade Debt [2] [https://lnkd.in/e-MkP49W](https://lnkd.in/e-MkP49W) First Brands Creditor Seeks Investigation of "Vanished" Cash [3] [https://lnkd.in/e-MWFVhs](https://lnkd.in/e-MWFVhs) First Brands’ Go-To Bank Jefferies Opens Books as Fallout Builds [4] [https://lnkd.in/eQfcRXFi](https://lnkd.in/eQfcRXFi) BlackRock Seeks Cash From Jefferies Fund Exposed to First Brands ## Why Rate Cuts Won't Fix a Labor Supply Shock - URL: https://sharifhsn.dev/blog/linkedin-2025-09-09-why-rate-cuts-won-t-fix-a-labor-supply-shock/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-09-09-why-rate-cuts-won-t-fix-a-labor-supply-shock.json - Description: Interest rate cuts won't relieve a labor supply shock: with lessons from energy and housing crises. - Date: 2025-09-09 - Exact published timestamp: 2025-09-09T19:37:25.419Z - Topics: Monetary Policy, Labor Markets - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7371265494668156929/ Interest rate cuts won't relieve a labor supply shock: with lessons from energy and housing crises. After today's stunning job revision of -911,000 jobs for March 2025, the stock market rallied. The S&P index is hitting all time highs, as it has been over the past two weeks. That might seem like a contradiction, but it's because the BLS data is used to guide the decisions of the Federal Open Market Committee in raising or lowering the interest rate. A bad jobs report will push the FOMC into lowering rates. Lower interest rates make it cheaper to borrow money across all sectors of the economy. This makes companies more likely to expand and therefore hire more workers. That's the thinking on how rate cuts counteract a bad labor market. But there's a problem with how this applies to the current macro environment. Think like a company. Just because it's cheaper to borrow money doesn't it mean make sense for to take on new ventures. If there's nobody to hire, then it doesn't matter how much money you have. Monetary policy is fundamentally a demand-side policy. Retiring boomers, decreased immigration, and general uncertainty over federal policy have made companies cautious in taking risks dependent on labor. The same situation happened in the 1970s during the energy crisis. When geopolitical instability caused supply shocks in oil, the US economy entered "stagflation", where both inflation and unemployment were increasing. The only solution was the "Volcker shock", when Fed Reserve head Paul Volcker sharply rose interest rates from 11.2% to 20% in order to crush the runaway inflation with a recession. And last year, despite rate cuts of a total 1%, interest rates on mortgages remained high, increasing from a high 6% to 7.1%. Typically, mortgage rates decrease when the Fed Funds rate does, as lenders try to attract new business with lower rates while existing homeowners refinance at lower rates. But the supply constraints caused by slowdowns in building homes meant that even if it was cheap to buy a home, people wouldn't be able to find one. With the increased media scrutiny on the Federal Reserve, along with the president's statements about them, their decisions matter more than ever. On September 17, we will see whether they cave to these expectations. [https://lnkd.in/gjBtmekN](https://lnkd.in/gjBtmekN) ## Oil Prices and the Energy Provisions in the One Big Beautiful Bill - URL: https://sharifhsn.dev/blog/linkedin-2025-07-03-energy-provisions-one-big-beautiful-bill/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-07-03-energy-provisions-one-big-beautiful-bill.json - Description: Great video about oil prices from experts in commodities. - Date: 2025-07-03 - Exact published timestamp: 2025-07-03T18:00:26.040Z - Topics: Energy Policy, Fiscal Policy - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7346598711562579969/ Great video about oil prices from experts in commodities. I think Wasif Latif's point about the price being right for oil is really significant for the implications of the energy provisions of the One Big Beautiful Bill set to pass in the House soon. Subsidies make sense for a growing innovative industry like solar, not so much for oil where geopolitical tail risk surrounding a massive supply is the biggest concern. Oil companies are smart. They're not going to just jump up and bite at oil subsidies and expand infrastructure when they know that the supply and demand factors don't favor them. ## Neural Networks Aren't Impressive on Their Own - URL: https://sharifhsn.dev/blog/linkedin-2025-06-07-neural-networks-aren-t-impressive-on-their-own/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-06-07-neural-networks-aren-t-impressive-on-their-own.json - Description: Neural networks AREN'T impressive... - Date: 2025-06-07 - Exact published timestamp: 2025-06-07T19:26:59.305Z - Topics: AI, Machine Learning - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7337198409189150720/ Neural networks AREN'T impressive... in the context of data analysis, by themselves. With the explosion in discourse about AI, traditional neural networks have also gained cultural prestige. LLMs do use neural networks as their underlying architecture, as do other forms of machine learning. At first blush, neural networks can accomplish some pretty cool results. For any input and output, neural networks can create a mapping that works effectively. The obvious problem is that without mitigation, neural networks will massively overfit the training data, and therefore be useless when applied to new data. But even when neural networks successfully fit test data, they are limited in what they can do. Sure, if your only goal is to predict outputs from an input, neural networks can do that for you. But neural networks (and nonparametric models in general) don't give a lot of general insights into relationships in your data. If you tried to approximate the relationship with a polynomial, you would get an ugly, long expression where the numbers don't mean anything. Compare to linear regression. You get a simple number of parameters on the same order of your features, and each number directly expresses the relationship of the feature with the label. Obviously, linear regression isn't feasible for expressing nonlinear relationships, which are common in data. And there's no parametric model that fits all types of data. That's kind of the point, though. By finding a model that fits your data, you're inherently saying something about the structure of the relationships in your data, which is often useful information. Neural networks are the lowest common denominator. You can put something in, and get something out, but for the real questions about your data, there's only a black box. They can be useful as a starting point. But they can't be the end of your data analysis. What do you think? Are neural networks useful in data analysis? ## The Case for Structured Errors in Rust - URL: https://sharifhsn.dev/blog/linkedin-2025-05-28-the-case-for-structured-errors-in-rust/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-05-28-the-case-for-structured-errors-in-rust.json - Description: I found this article on structured errors insightful. Although it is focused on Rust, the explanations of the direct benefits of encoding information about the failure surface of c… - Date: 2025-05-28 - Exact published timestamp: 2025-05-28T18:20:44.156Z - Topics: Rust, Software Engineering - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7333557857549783040/ I found this article on structured errors insightful. Although it is focused on Rust, the explanations of the direct benefits of encoding information about the failure surface of code are succinct and persuasive. "Custom error types allow you to see all potential failure modes of a function at a glance, without (recursively) inspecting its implementation or maintaining fragile hand-written docs that can’t be fully trusted anyway. In a code review, you can easily notice when some error variant doesn’t make sense and should be handled locally, or comes from an action that shouldn’t be performed at all. Interfaces become more descriptive." [https://lnkd.in/esm4w9WF](https://lnkd.in/esm4w9WF) ## Rust's `subsecond` Crate Brings Hot Reloading - URL: https://sharifhsn.dev/blog/linkedin-2025-05-28-rust-s-subsecond-crate-brings-hot-reloading/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-05-28-rust-s-subsecond-crate-brings-hot-reloading.json - Description: Rust's subsecond crate is a game changer for application and game development in Rust! Its innovative hotpatching method enables hot reloading for select libraries. - Date: 2025-05-28 - Exact published timestamp: 2025-05-28T18:00:05.331Z - Topics: Rust, Developer Tools - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7333552661541117952/ Rust's `subsecond` crate is a game changer for application and game development in Rust! Its innovative hotpatching method enables hot reloading for select libraries. By far the biggest drawback of Rust is its compilation time. Although it has improved over the years, it is far slower than any other language. Even an extra wait of a few seconds can break developer flow. This is, unfortunately, by design. C/C++ are able to compile quickly because they are generally split up into source code and header files which describe the functions individually, so the compiler can compile them in parallel if they don't depend on each other. Rust, however, treats the entire package as the compilation unit. Library authors have to split up their code, often awkwardly, in order to achieve faster compilation times. Dioxus Labs have finally engineered the solution for their GUI library Dioxus. Their innovative `subsecond` crate compiles only the part of the code that has changed and patches the change into the existing binary, significantly reducing compilation time for small changes. This is common when designing UIs, when shifting elements and changing small details. Most significantly, `subsecond` is available on ALL major platforms, which has never been achieved by ANY hot patching system for a compiled language. The developers behind the game engine Bevy were able to integrate `subsecond` into Bevy and hot patch game development with 500 ms compilation times, basically as fast as you can press Ctrl + S! Currently, `subsecond` is too bespokely designed to be used widely throughout the ecosystem. However, the Dioxus team is actively working to improve `subsecond`, and hopefully it will become an option for developers working on large, complex libraries. `subsecond` hot reloading will be included in the upcoming release of Dioxus 0.7, and likely Bevy 0.17. The Dioxus team will be presenting technical details about `subsecond` at Seattle RustConf from September 2-5. Dioxus release post: [https://lnkd.in/eJfHXWJH](https://lnkd.in/eJfHXWJH) Example of Bevy running `subsecond`: [https://lnkd.in/e457vQX9](https://lnkd.in/e457vQX9) ## What a U.S. Credit Rating Downgrade Means - URL: https://sharifhsn.dev/blog/linkedin-2025-05-17-what-a-u-s-credit-rating-downgrade-means/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-05-17-what-a-u-s-credit-rating-downgrade-means.json - Description: Moody's has downgraded the US's credit rating from the best (Aaa) to second-best (Aa1), sending the markets into a buzz. But what does it mean to downgrade a credit rating? - Date: 2025-05-17 - Exact published timestamp: 2025-05-17T15:52:46.770Z - Topics: Fiscal Policy, Credit Ratings - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7329534356572508160/ Moody's has downgraded the US's credit rating from the best (Aaa) to second-best (Aa1), sending the markets into a buzz. But what does it mean to downgrade a credit rating? All bonds have a property called credit risk. This is the probability that the seller of the bond will default on their debt and not pay it back in full or on time. This goes for all kinds of debt, whether it's a corporate bond, a mortgage, or credit card debt. The quantification of credit risk is important for pricing financial instruments. The reason that the interest rate on a credit card is much higher than a mortgage is because lenders know that people are much more likely to default on credit card debt than mortgages, and want to be compensated more for taking on that risk. (A mortgage also has a lower interest rate because it's secured debt, but I'll ignore that for the purposes of this explanation) For corporate and sovereign (fancy word for national) bonds, this risk is expressed through credit ratings. Ratings agencies like Moody's, S&P, and Fitch look through financial statements, the macroeconomic outlook, and historical data to determine the credit risk of a debt issuer. They will then issue ratings based on their assessment, which can then be directly translated into an equivalent hazard rate (probability of default over a given period) for the purpose of pricing credit derivatives. This downgrade reflects Moody's belief that, as a result of mounting US debt, they have increased in credit risk, a belief shared by S&P and Fitch, who have both already downgraded the US in a similar way. A Saturday surprise like this will no doubt impact markets soon. Previous downgrades have led to yield spikes and market rallies. The 10-year yield is already going up, and we'll see how the markets will look on Monday. What do you think? Does the US deserve this downgrade? From Bloomberg: [https://lnkd.in/e9yQ9u2E](https://lnkd.in/e9yQ9u2E) ## Rust at 10: Why I've Followed the Language - URL: https://sharifhsn.dev/blog/linkedin-2025-05-16-rust-at-10-why-i-ve-followed-the-language/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-05-16-rust-at-10-why-i-ve-followed-the-language.json - Description: Rust turns 10! In honor of this momentous anniversary, I'd like to reflect on what Rust means to me, and why I have followed it for so long. - Date: 2025-05-16 - Exact published timestamp: 2025-05-16T17:38:11.094Z - Topics: Rust, Programming Languages - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7329198494844411905/ Rust turns 10! In honor of this momentous anniversary, I'd like to reflect on what Rust means to me, and why I have followed it for so long. I first heard about Rust in 2021, when I was in my third year of my computer science degree. At this point, I had taken and was taking multiple classes in C, like Computer Architecture, Systems Programming, and Operating Systems. In each of these classes, I had been immensely frustrated with the memory errors I would get and their inscrutability. To this day I have a PTSD reaction to the phrase "Shadow bytes around the buggy address". Rust seemed like the perfect solution. Memory safety with all the same low-level interaction? Sign me up! But it became much more than that to me. I noticed that algebraic data types, which I had learned in OCaml in my PrinProg class, were present in Rust, and I was surprised by how useful they were in modeling state, having only used them in a toy context. From generics to the module system, everything about Rust seemed to have learned from the lessons of previous programming languages and improved on them. Rust was also constantly evolving, which made it exciting to follow. Every six weeks there was a new release, which would invariably have some new feature that people had been requesting. Core libraries like async, error handling, and time objects were shifting in prominence, constantly adding new features. Large projects like the game engine bevy and the web server axum would wow the community on every new version. At the same time, Rust was being slowly integrated into the business world. The first big news was a blog post from Discord describing how they used Rust to greatly optimize a service they had previously relied on Go for. Microsoft, Google, and Amazon all began to trickle out announcements of how they used Rust to optimize heavily used core functionality. The biggest deal by far was the announcement of Rust for Linux, where the Linux kernel would, for the first time, allow a language other than C (or Assembly) into the source code. Although it has pretty much only been used for drivers so far, the symbolism of being included in Linux cannot be overstated. Ten years later, the third-party ecosystem has stabilized significantly, and more and more companies are considering Rust as the developer base grows. Only time will tell where Rust goes from here, but I can safely say that we'll only hear more about it. More reflections on Rust from well-known community members: ★ Graydon Hoare, creator of Rust: [https://lnkd.in/eagHb3Xc](https://lnkd.in/eagHb3Xc) ★ Steve Klabnik, a founding member: [https://lnkd.in/esJy_PZV](https://lnkd.in/esJy_PZV) ★ Niko Matsakis, compiler theorizer: [https://lnkd.in/eKQytR2Z](https://lnkd.in/eKQytR2Z) ## Copulas and Portfolio Risk - URL: https://sharifhsn.dev/blog/computational-methods-week-14/ - Structured data: https://sharifhsn.dev/api/posts/computational-methods-week-14.json - Description: Ending early today because the Pacers are playing tonight - Date: 2025-04-29 - Exact published timestamp: 2025-04-29 - Topics: Computational Methods, Copulas, Value at Risk, Expected Shortfall - Categories: Computational Methods - Source: Computational Methods in Quantitative Finance - Source URL: None Ending early today because the Pacers are playing tonight ## Copulas Why do we need this? First let’s talk about: Generally speaking, the **Basel 2**, Basel 3, and current Basel 3.5, all these rules and regulations require that banks assess two things. Any unit that owns funds that don’t belong to them (banks, basically), they need to assess risk. What exactly does that mean? The class 535, 635, and 636 are all about this. 636 will be offered this fall. At the core, regulators don’t understand any of it. What they want to know is, if the economy goes to crap, how much will you lose? So how do you come up with this number? There’s two basic numbers: **VaR**, or Value at Risk, and **ES** Expected Shortfall One of my former students is head of globalr isk at Barclays, and every day he has to come up with these numbers. Basically he aggregates the positions of every desk at Barclays. Derivatives, asset management, whatever. Each has a risk associated with it. VaR is supposed to be a quantile. Imagine you have a distribution for the value of your assets tomorrow. It’s actually in 10 days that Basel demands.Obviously, I know the value of assets today, by marking to market. I project in ten days what are the possible things that could happen to all my divisions. My derivatives desk has a distribution, another desk has another one, so I have a whole bunch of random variables. Now I need to combine them. I have tons of distributions, how do I put them together? For example, 10% of my bank is invested in derivatives. So I should take the risk from that and multiply by 10%, right? Wrong. Because things don’t move independently. We could do such aggregation only then. If the derivatives desk goes down because Trump said something dumb, th equities desk will also go down, and so will bonds. Things are related. That’s the purpose of copulas. These are very much used in risk management. VaR can only be calculated for a 1D r.v. You need to have a typical marginal distribution. Let’s say you get this distribution, of values with probabilities. Let’s say I look at the value where the probability is 0.01. If I look at the value I have now, and subtract the value there, that resulting value is the loss. The amount of money I will lose if this thing happens, which is worst case scenario, with a 1% chance of happening. This particular one is VaR\_99. You can obviously calculate different percentiles. ### Expected Shortfall If you are in this region, you are going to lose money. If my estimation is correct, then this thing is going to happen. If I look at the next 100 days, and do this distribution every single day, then I expect to see one of these events every 100 days. But if it does happen, exactly how much money will I use? I don’t know which point I’m in, just the region. So to determine that value, you calculate ES. ES is nothing more than $$\\mathbb{E}\[\\text{Loss}|X \< \\text{VaR}\]$$ Basically it’s just the expected value of the random variable when it’s less than VaR. This is important, there’s an entire area of risk management dedicated to this. The trader cannot take risks because the risk manager won’t let them, because of this stuff. You can’t take crazy risks. ### Marginal vs Joint Distribution So what is the problem? This is calculated on a daily thing. A natural thing to do is to look back 100 days, have a historical look at each desk. That gives me distributions that are reasonable for each of those rvs. If I have a 1D rv, with 100 observations, I can make a pretty accurate histogram. For a 2D rv, I suddenly need a lot more observations. And for more, it’s impossible. So although you can estimate the distribution of the marginals very easily, you can’t do that for the joint. You would need mountains of data, which is not possible. And the problem is if you go too far in the past, you end up with irrelevant data that is nonstationary. The question is, can we do that? You can do this in one case, and only one case. And that’s if X\_1 … X\_d are independent. Then the joint distribution is the product of the marginals. This holds for both PDF and CDF. However, irl, nothing is independent. In the 1940s, there’s an old theorem by **Sklar**: *Let X\_1…X\_d random variables with CDFs F\_1…F\_d.* (These are one dimensional CDFs) *Let* $$F(x\_1, \\ldots x\_d) \= \\mathbb{P}(X \\leq x\_1, \\ldots, x\_d \\leq x\_d)$$ There exists a function C which follows $$\\mathbb{R}^d \\rightarrow \[0, 1\]$$ That this function $$C(F\_1(x\_1), F\_2(x\_2), \\ldots F\_d(x\_d) \= F(x\_1, \\ldots x\_d)$$ In the case of independent rvs, then it’s just the product. But maybe we have to do something further. Then if you have the joint pdf f\_1(x\_1 … x\_d), then you also have a similar $$f(x\_1, \\ldots, x\_d) \= c(F\_1(x\_1), \\ldots F\_d(x\_d)) f\_1(x\_1) \\ldots f\_d(x\_d)$$ And this is just how derivatives work, this is the chain rule, where $$c \= \\frac{d}{dx\_1} \\frac{d}{dx\_2} \\ldots \\frac{d}{dx\_d} C$$ There’s a second property that C is unique on \[0, 1\]^d. Basically, the function c connects the marginals, which are probabilities \[0, 1\]. So you limit yourself to the space of distributions, which are the $$\\mathbb{R} \\rightarrow \[0, 1\]$$, then we are unique. This is pretty powerful. ## Issues with Copulas There is a non obvious issue: **time stationarity**. This is not even mentioned in most texts. When you estimate this, you need to estimate the marginals. You need a mountain of data to have these two things. You calculate the copula which connects them. Let’s say you have 2001-2006 and estimate the copula during that period. In order to make forecasts about the relationship in the future, you’re assuming that the relationship stays the same. This is the reason for the 2008 financial crisis. The CDOs are estimated assuming the functions doesn’t change, which is pretty stupid. Things don’t seem correlated right now, but they were, even though the historical data didn’t show it. The second big issue: what EXACTLY is this copula? We know this function exists, but we don’t know how to get to it. It’s like God. We know God exists, but if we go by Christian religion it’s this god, if it’s Islam it’s that god, etc. ## Copula Families I’ll say that a copula function, sorta, looks like this. The only quantity I know is that it’s defined from \[0, 1\] to \[0, 1\]. So if I have a function that does this, I should be okay. If you take 1D CDF, then what does it do? This is a function $$F: \\mathbb{R} \\rightarrow \[0, 1\]$$. It takes a random variable, and gives it a probability. What about $$F^{-1}$$, the quantile function? This is defined $$F^{-1}: \[0, 1\] \\rightarrow \\mathbb{R}$$ All copulas are based on this. They use the cdf and the inverse. The most popular one, we have already seen. The simplest one is called ## Gaussian Copula The whole thing is that the function needs to associate $$\\mathbb{R}^d \\rightarrow \[0, 1\]$$ We have the multivariate normal function $$\\Phi\_\\Sigma$$ CDF.with mean 0 and covariance matrix Σ. The cov matrix is positive, semidefinite, and symmetric. $$\\Phi\_\\Sigma (u\_1 \\ldots u\_d) \= \\int\_{-\\infty}^{x\_1} \\int\_{-\\infty}^{x\_d} \\frac{1}{\\sqrt{(2\\pi)^d \\text{det} \\Sigma}}e^{-\\tfrac{1}{2}u^T \\Sigma^{-1}u} du\_1, \\ldots du\_d$$ u is a line vector This has to be between 0 and 1\. But this is defined for R. So I’m going to write the gaussian copulas as $$C^{\\text{Gaussian}} (x\_1 \\ldots x\_d) \= \\Phi\_\\Sigma (\\Phi^{-1}(x\_1), \\ldots , \\Phi^{-1}(x\_d))$$ This is really complicated. The Fs are the xs, you plug in the marginals, and that gives you the copula function. The important concepts are the multivariate Gaussian with given sigma, the Φ-1 inverse cdf, and you plug in the cdfs of the marginals. You look at your data, obtain the marginals, and obtain the joint. And fit the function to the joint. The only thing to estimate here is Σ. ## Student t Copula There is another thing called the Student t copula. Remember the Cholesky decomposition? It’s basically used for this. You use the data, you use the numbers you get, and calculate the joint distribution. The correlated numbers you plug into the copulas, and you get the thing correlated. Fancy way of doing the homework. It’s an extension that uses the fact that $$t \= \\frac{Normal}{\\Chi^2}$$ You obtain χ^2 by squaring the normal. This is useful. Gaussian copula is nice and easy to work with. ## Archimedean Copulas This is an entire class. We have Gaussian, and student t which inherits from Gaussian processes. (By the way, if you want to know more about Gaussian processes, take Florescu’s class). The whole point is that you’re going with cdfs, and inverse cdfs. Remember a copula goes from \[0,1\]^d to \[0,1\]. It is defined like this $$C(u\_1, \\ldots u\_d) \= \\Psi^{-1}(\\Psi(u\_1) \+ \\ldots \+ \\ldots \\Psi(u\_d))$$ where $$\\Psi: \[0, 1\] \\rightarrow \[0, \\infty)$$ and $$\\Psi(1) \= 0$$ This Ψ function is kinda like inverse cdf. Importantly, Ψ^-1 is monotonic. And it follows that Ψ^-1(0) \= 1 Ψ goes from \[0, 1\] to ∞, and the inverse goes the other way. And there’s a property that Ψ^-1 starts at 1, and keeps going down over to ∞. ## Product Copula You take the $$\\Psi(u) \= \-\\ln u$$ $$-\\ln u \= y$$ $$u \= e^{-y}$$ So that’s the inverse $$\\Psi^{-1}(y) \= e^{-y}$$ Remember this starts at 1 at x \= 0, then decreases towards ∞. The archimedean copula function is the sum of Ψ. So what is the copula in this case? $$C(u\_1, \\ldots u\_d) \= \\Psi^{-1}(\\Psi(u\_1) \+ \\ldots \+ \\Psi(u\_d))$$ $$= e^{-(-\\ln u\_1 \- \\ln u\_2 \- \\ldots \- \\ln u\_d)} \= e^{\\ln(u\_1 \\ldots u\_d)}$$ which ends up being $$= u\_1 u\_2 \\ldots u\_d$$ You take each cdf and multiply them, which is numbers between 0 and 1\. It’s kind of useless because it doesn’t have a parameter, so you can’t fit it. ## Clayton These are the most used ones $$\\Psi(u) \= \\frac{1}{\\theta} (u^{-\\theta} \- 1)$$ Basically you construct these, by picking something. Someone came up with a function such that the inverse looks like this, then they pick θ such that they have a parameter to fit something, where Ψ(1) \= 0\. They actually are simple, and then they just fit the properties. SOmoene tried it, now everyone uses it. The inverse is $$\\Psi^{-1} \= (\\theta y \+ 1)^{-\\tfrac{1}{\\theta}} $$ $$C^{\\text{Clayton}} (u\_1 \\ldots u\_d) \= (\\theta(\\Psi(u\_1) \+ \\ldots \+ \\Psi(u\_d)) \+ 1)^{\\tfrac{1}{\\theta}}$$ Then just plug in the Ψ term. $$= (\\theta(\\sum\_{i=1}^d \\frac{1}{\\theta}(u\_i^{-\\theta} \- 1)) \+ 1)^{\\tfrac{1}{\\theta}}$$ $$ \= (\\sum\_{i=1}^d u\_i^{-\\theta} \- d \+ 1)^{\\tfrac{1}{\\theta}}$$ This formula IS NOT ONLINE, only for bivariate copula. All the formulas online are bivariate I have four more pages of crap for Frank and Gumbel. But… ## Frank This is a generator $$\\Psi(u) \= \-\\log(\\frac{e^{-\\theta u} \- 1}{e^{-\\theta} \- 1})$$ ## Gumbel $$\\Psi(u) \= (-\\log u)^\\theta$$ And if you understand the principles, you can come upw ith your own. you can notice that you have logs everywhere, because it’s very nice because when you exponentiate it disappears. But it could be something like cosine\! The only requirement is that the inverse function is convex. The theta is the parameter fit. The way you fit theta, the relationship is $$C(F\_1(x\_1), \\ldots F\_d(x\_d)) \= F(x\_1 \\ldots x\_d)$$ You’re just inventing something to fit the copula, something close. It’s a high barrier of entry because the formulas seem complicated. ## Why Do We Care We need to estimate the wealth and know the risks, and come up with a number. But how do we actually use this stuff? I used to do a lecture on generating random variables. Now I don’t. If you’re interested in more detail, read the Handbook of Probability. Chapter 6 proves this. If you take a random vairalbe and plug it into a cdf, it generates a uniform random variable. This tells you how to generate any rv you want. All you have to do is the inverse cdf method. If $$U \\sim \\text{Uniform}(0, 1)$$, and you can get the inverse cdf $$F^{-1}(U)$$ has the same distribution as X. If you can calculate the CDF, and the inverse CDF, you can generate random uniform, plug it into it, and then you get a random. The problem is the normal CDF, because the normal CDF is not invertible. But a lot of other rvs work. The problem is that it only works for 1D rvs. That’s why this copula is pretty useful. Recall that the joint distribution cdf is equivalent to the copula of the marginal cdfs. The methodology will work for the marginal cdfs, so you can apply the copulas to get to the joint cdf. ### Algorithm First you estimate C. 1: First you generate many uniforms iid u\_1… u\_d 2: For each such point, you calculate the C(u\_1, … u\_d) for all generated u\_is. 3: Calculate marginals densities 4: I now have all these probabilities that I generate x\_1… x\_d. My x\_i cdf can be used to generate numbers, I have d distributions, I have a number between 0 and 1, and extract the corresponding x\_i, which gives me a generate number. That will calculate not only VaR but also ES with respect to these distributions. There could be an easier version of this, which is what Marina will do. This is just how I would do it. ## Counterparty Risk-Weighted Assets - URL: https://sharifhsn.dev/blog/counterparty-risk-weighted-assets/ - Structured data: https://sharifhsn.dev/api/posts/counterparty-risk-weighted-assets.json - Description: The risk-weighted-assets notes translate counterparty exposure into a capital calculation. The exposure-at-default measure is combined with a risk weight or a model-based capital r… - Date: 2025-04-28 - Exact published timestamp: 2025-04-28 - Topics: FX, Counterparty Risk, Risk-Weighted Assets - Categories: FX - Source: FE-635 \| Risk Engineering - Source URL: None The risk-weighted-assets notes translate counterparty exposure into a capital calculation. The exposure-at-default measure is combined with a risk weight or a model-based capital requirement, with maturity, collateral, netting, and recovery assumptions affecting the result. The important control is consistency. Current exposure, potential future exposure, and the counterparty's default probability describe different parts of the problem; substituting one for another can make a capital number look precise while measuring the wrong risk. FE-635 treats the calculation as a modeling and governance exercise: document the inputs, preserve the quote and currency conventions, and make the scenario behind the capital estimate reproducible. ## Multi-Name Credit Risk - URL: https://sharifhsn.dev/blog/advanced-derivatives-week-13/ - Structured data: https://sharifhsn.dev/api/posts/advanced-derivatives-week-13.json - Description: This is going to be a basic introduction, this is a complex topic to cover comprehensively. - Date: 2025-04-24 - Exact published timestamp: 2025-04-24 - Topics: Credit, Copulas, Dependence Modelling, Multi-Name Credit Risk - Categories: Credit - Source: Advanced Derivatives - Source URL: None ## Copulas This is going to be a basic introduction, this is a complex topic to cover comprehensively. ## Limitations of Multi-Name Latent Variables First of all, let’s start with some considerations with respect to the previous model we discussed, some of its limitations. The Multi-Name latent variable model is single factor. There is one variable which is seen by the other credits. They each have their own idiosyncratic component. In terms of properties, simulations, that’s it. We looked at conditional hazard rate and calculation of the portfolio of loss. Last assignment is related to this particular topic. We can model the correlation, we can simulate, we can calibrate, the time-dependent threshold, based on the hazard rate and survival probabilities. How can we extend this model? There’s **limitations on the one-factor correlation structure**. In one-factor, we have a number of correlation parameters N\_c. Recall that the correlation between two credits i and j is beta\_ij. We have the constraint that the absolute value of each correlation parameter is less than or equal to 1\. The entire correlation matrix has \\(\\frac{N_c(N_c-1)}{2}\\) correlation parameters, which is a lot. The question we can ask: Does one factor structure prevent the modeling of groups, or **sectors**? The model may be insufficient. We can look at some examples. Consider a simple portfolio of 4 credits grouped in two sectors. How can we represent such groups? We could take beta\_1 \= beta\_2 \= beta\_a. These two stocks will have the same exposure to the market variable z, and vice versa for the other sector beta\_3 \= beta\_4 \= beta\_b. In a single factor model, the correlation is the product of two correlation parameters for two credits. $$ C = \\begin{bmatrix} 1 & \\beta_a^2 & \\beta_a \\beta_b & \\beta_a \\beta_b \\ \\beta_a^2 & 1 & \\beta_a \\beta_b & \\beta_a \\beta_b \\ \\beta_a \\beta_b & \\beta_a \\beta_b 1 & \\beta_b^2 \\ \\beta_a \\beta_b & \\beta_a \\beta_b & \\beta_b^2 & 1 \\end{bmatrix} $$ If you split this into four square matrices, then you’ll see the sub-correlation matrix for each group. The correlation between the two groups. Let’s assume that beta\_a \> beta\_b \> 0\. One is more exposed to the market than the other. This also implies that $$ c = \\beta_a^2 \\geq \\beta_a \\beta_b \\geq \\beta_b^2 $$ So if you take a numerical example, SEE NOTES PAGE 3 Intersector correlation is much higher, however. We don’t want that. 40% \> 25% is a big problem. Sector B is not fully captured. So we can try to extend the model, make it more complex, and accommodate this particular scenario. In this case, we can consider a **two-factor** version of the model. For some credit i, we have $$ A_i = \\beta_{1i} Z_1 + \\beta_{2i} Z_2 + \\sqrt{1 - \\beta_{1i}^2 - \\beta_{2i}^2} \\epsilon_i $$ $$ A_j = \\beta_{1j} Z_1 + \\beta_{2j} Z_2 + \\sqrt{1 - \\beta_{1j}^2 - \\beta_{2j}^2} \\epsilon_j $$ Similar to the previous formulation, but with one more factor. The linear combination of standard normal is standard normal, so that doesn’t change. The latent variable A\_j corresponds still to the time-dependent threshold for each credit. We can calculate the correlation between these assets in a similar kind of matrix. For each element: $$ c_{ij} = \\beta_{1i}\\beta_{1j} + \\beta_{2i}\\beta_{2j} $$ We can revisit the matrix in the following case. In sector A, the credits have a first factor weight and second factor weight beta\_1a and beta\_2a. Sector B is similar with beta\_1b and beta\_2b This is more generic, you would say this group has correlation to these factors. We’ll assume the first credits are from sector A. That would be the correlation between credit a in sector a and b in sector b $$ c = \\begin{bmatrix} 1 & \\beta_{1a}^2 + \\beta_{2a}^2 & \\beta_{1a} \\beta_{1b} + \\beta_{2a}\\beta_{2b} & \\beta_{1a} \\beta_{1b} + \\beta_{2a}\\beta_{2b} \\ \\beta_{1a}^2 + \\beta_{2a}^2 & 1 & \\beta_{1a} \\beta_{1b} + \\beta_{2a}\\beta_{2b} & \\beta_{1a} \\beta_{1b} + \\beta_{2a}\\beta_{2b} \\end{bmatrix} $$ (Bottom half is same as top half, mirrored) Example on notes page 6 The credits in two groups will have different correlations on the second factor. In general, M-factors are needed to model M-sector portfolio. Three sectors? Three factors. The simulation is very simple, but instead of generating just a Z\_1, we generate Z\_1 and Z\_2. It becomes a little more complex, with the number of parameters, compared to a single factor. We go from n(n-1)/2, to a lot more. If you want to implement these, you need to calibrate the exposure to the parameters. So the complexity becomes prohibitive, and that’s a limitation. It has some applicability, but in general, the more typical way is through copula functions. ## Copulas (Finally) You can model default time dependencies, correlations for default times. We also will have measures of dependency. We will look at main copulas for credit modeling, and Monte Carlo pricing of correlated products with copulas. We’ll look at some properties, and dependency limits, and some relations between copulas and the model we just discussed. And other ways of modeling dependence than correlation. If we have time, we’ll have some code to visualize copulas. Next week, I have a section for estimation of parameters of copulas, more complex. But for the final exam, copulas will not be included. The latent variable model can also be expressed as a Gaussian copula, so they are related. **Definition: An N-dimensional copula function C is a multivariate cumulative distribution function with N uniform marginals with probabilities u\_1, u\_2, … u\_N.** This is just the multivariate cdf. One of the characteristics is that it has uniform marginals. So we can say that thIS IS THE JOINT PROBABILITY $$ \\mathbb{P}(\\hat{U_1} \\leq u_1, \\hat{U_2} \\leq u_2 \\ldots \\hat{U_N} \\leq u_N) = C(u_1, u_2 \\ldots u_N) $$ This is the definition of such copulas where $$ \\hat{U_i} \\sim \\text{ Uniform}(0, 1) $$ ## Dependent Structure of Default Times of N Credits The idea here is to express this in terms of copulas. Consider random variable u\_i at time t\_i $$ u_i(t_i) = 1 - Q_i(t_i) $$ This is the probability of credit i defaulting before t\_i. Remember, this u\_1 is a probability which can take values. So the copula function is the same $$ C(u_1(t_1), u_2(t_2), \\ldots u_N(t_N)) = \\mathbb{P}(\\tau_i \\leq t_i, \\tau_2 \\leq t_2, \\ldots \\tau_N \\leq t_N) $$ where tau is the time of default. In this copula, it’s expressed in terms of the default probabilities. So this is the “default copula”. We have this default copula, so we can also talk about the survival copula. There are some properties of copulas. * Because cdf increases always, we know that if any name’s default probability increases, then the total joint probability of default will increase. * We can get the marginal distribution of u\_k if u\_i \= 1 for all i \!= k. In order to get the marginal of any of these names, you can take \\(C(1, 1, 1, \\ldots u_k, 1, \\ldots 1) = u_k\\) * When is the copula equal to 0? If any of them have zero probability of default, riskless like Treasury bonds, then the copula will be 0\. * The dimensionality of the copula can be reduced from N to N \- 1 dimensions by setting any u\_i \= 1\. This is **very important** because it’s still a copula if N \> 2\. N \= 2 is useful for examples. Most textbooks describe properties in 2 dimensions. But then it’s easy to extend into N dimensions, because you can set these u\_is to 1 and collapse it. ## Dependency Limits Latent variable model has **dependence limits**. We kind of look to see when we have positive and negative dependence of names in the portfolio. This is not equal to the correlations, but we can think a little bit in that way. When we have an expression for a copula, (the joint probability distribution), it’s good to look at these properties to see the maximum and minimum dependence, and then how to measure them, based on differences dependency metrics. ## Independence The simplest is independence. The expression for such an independence copula is $$ C(u_1, u_2, \\ldots u_N) = u_1 u_2 \\ldots u_N = \\prod_{i=1}^N u_i $$ If they’re independent, the joint probability distribution is the product of the marginals. ## Perfect Positive Dependence ## Perfect Negative Dependence We can then state the **Fréchet-Hoeffding Theorem:** For any copula \\(C: [0, 1]^N \\rightarrow [0, 1]\\) and any realization of \\((u_1, \\ldots u_N) \\in [0, 1]^N\\), the following bounds hold: $$ w(u_1, u_2 \\ldots u_N) \\leq C(u_1, u_2 \\dots u_N) \\leq M(u_1, \\ldots u_N) $$ Basically this theorem says if you have a copula with a particular realization, this copula is bounded above and below by two functions M and w. w the lower Fréchet-Hoeffding bound is $$ w(u_1, \\ldots u_N) = (1 - N + \\sum_{i=1}^N u_i)_+ $$ and upper bound is $$ M(u_1, \\ldots u_N) = \\min(u_1, \\ldots u_N) $$ In the bivariate case where we only have two variables u and v, we have $$ \\max(u + v - 1, 0) \\leq C(u, v) \\leq \\min(u, v) $$ Upper bound is perfect positive, lower bound is perfect negative. This result holds for **any copula**. Ideally we would like to estimate this copula with some data. We need to make a choice for this expression, and estimate those parameters. But the question is, this is a representation of the original joint probability distribution, but is this a unique representation? Can we have different copulas for this? We have one more theorem, **Sklar’s Theorem**. We know by definition, the copula is a multivariate distribution function. Any such function can be written as a copula, and its representation is unique, the marginal distributions are continuous. If you have discrete distributions, then uniqueness is not guaranteed. Let’s consider more generally, any N-dimensional distribution function H with marginal distributions F\_1, F\_2, F\_N for random variables x\_1, x\_2, … x\_N. H is the joint probability distribution. Because this is multi-dimensional, this will correspond to $$ H(x_1, x_2, \\ldots x_N) = \\mathbb{P}(X_1 \\leq x_1, X_2 \\leq x_2 \\ldots, X_N \\leq x_N) $$ We can write this as the probability that the marginal function F\_1 $$ = \\mathbb{P}(F_1(X_1) \\leq F_1(x_1), F_2(X_2) \\leq F_2(x_2), \\ldots F_N(X_N) \\leq F_N(x_N)) $$ We basically want to transform this to get a copula function, to get here from a generic N-dimensional function. Let \\(\\hat{U_i} = F_i(X_i)\\) and the realization \\(u_i = F_i(x_i)\\) The marginal will take values between 0 and 1, it’s a cdf. The original function $$ H(x_1, x_2 \\ldots x_N) = \\mathbb{P}(\\hat{U_1} \\leq u_1, \\ldots \\hat{U_N} \\leq u_N) $$ $$ H(x_1, x_2 \\ldots x_N) = C(F_1(x_1), F_2(x_2) \\ldots F_N(x_n)) $$ We know that C is a unique copula if these Fs are continuous. And in reverse, $$ C(u_1, u_2 \\ldots u_N) = H(F_1^{-1}(u_1), F^{-1}_2 (u_2), \\ldots F^{-1}_N (u_N)) $$ The relationship between these two is that the copula is going to separate the choice of the marginals from the choice of the dependency structures. It decouples these things. We can take any function and its copula representation, and calibrate it to the names of the portfolio independent of the dependency structure based on the way I want to assign dependency to this portfolio. We might want to add some dynamics to this dependencies, the difference between credit default so we can calculate the conditional survival/default probability. We could impose this by choosing a different dependency structure. ## Alternative to Dependence Structure of Default Times We discussed about the default copula. We have 1 \- the default probabilities, which is the survival copula $$ \\hat{C}(1 - u_1(t_1), 1 - u_2(t_2), \\ldots, 1 - u_N(t_N)) = \\mathbb{P}(\\tau_1 > t_1, \\tau_2 > t_2 \\ldots \\tau_N > t_N) $$ Typically survival copula has this little hat 🙂 We will look at the bivariate case, but in general we have a relationship between the survival and default copula. $$ \\hat{C}(u, v) = u + v - 1 + C(1 - u, 1 - v) $$ These are some kind of definitions. We’ll make a short connection with the previous model. The latent variable model is a Gaussian copula model. If you consider this model, how would we express this? [The displayed equation following this prompt was not recoverable from the exported note.] That’s the probability that the time dependent threshold of credit i is smaller than the time dependent threshold We have a bivariate normal distribution: \\(\\Phi_{2, \\rho}\\) So how do you write this in terms of default? See page 18 of Notes It has the same representation of bivariate normal. This is a Gaussian copula with marginals \\(F_i(x_i) = \\Phi(x_i)\\) Therefore the Gaussian bi-variate CDF is notated by $$ C_\\rho^{GC} = \\Phi_2[\\Phi^{-1}(u_1), \\Phi^{-1}(u_2)] $$ And then we end up with an analytical expression. The survival copula is the same, just with 1 \- u\_1 and 1 \- u\_2. And because of the properties of the cdf, you can just make it the negative inverse cdf instead of 1 \- cdf. And actually, since there are two negatives, they cancel out. So there is symmetry: **the survival and default copulas of the Gaussian bi-variate are the same**. Because of this property, this particular distribution does not have tail dependence. To understand this, we will look at various dependency measures. Other distributions will model this more accurately. ## Measuring Dependence An important property of this is dependence, so the question is how do you measure? There are many ways. The most popular way is **Pearson Linear Correlation**. $$ \\rho_P = \\mathbb{E}[XY] = \\frac{\\mathbb{E}[X]\\cdot\\mathbb{E}[Y]}{\\sqrt{\\mathbb{E}[X^2] - (\\mathbb{E}[X])^2)}\\sqrt{\\mathbb{E}[Y^2] - (\\mathbb{E}[Y])^2)}} $$ Advantages: easier to understand. Everyone understands correlation. It is also invariant under linear transformation. Say we change the variables by scaling and shifting (multiplying and adding), it doesn’t affect the correlation. Another advantage is that if the marginals are Gaussian and the correlation is 0, then we have independence. However, there is a big disadvantage: linear correlation being 0 DOES NOT imply independence in general Sometimes the correlation is abused. It is a measure of linear dependence between variables. But just because there’s no linear dependence, doesn’t mean there are other kinds of dependence. A quick counterexample is on page 21 of the notes. Let’s say we have an experiment and we get five points, in y \= x^2. And we want to determine the correlation. Covariance is the same thing as the top of Pearson linear correlation. Expectation of x and x^3 are both 0, so everything is 0\. But in this case, they are strictly dependent, they just have no linear dependency. ## **Rank Correlation** If we have x, y random variables and draw n pairs (X\_i, Y\_i). Define R\_i as the rank of X\_i. That will tell us the order/rank of the variables in our realization, in terms of smallest to largest. Basically, we can take a sorted vec, and map the indices of the sorted vec to the elements. If S\_i is the rank of Y\_i, then the average rankings you can express rank correlation in terms of prevalence of concordant and discordant pairs. If we have two realizations (x\_1, y\_1) and (x\_2, y\_2), the pairs are concordant if \\((x_1 - x_2)(y_1 - y_2) > 0\\) The pairs are numeric values. It’s a product of two numbers, if it has to be greater than 0, either they are both positive or both negative. Discordant is if it’s negative. Example on page 24 of the notes If we measured the rank correlation, it would be 100%. If we look at these data points. How do we measure this? There are two ways. ### Kendall’s Tau $$ \\tau = \\frac{c - d}{c + d} $$ where c is the number of concordant pairs, and d is the number of discordant pairs. Then the total number of pairs is c \+ d, which n choose 2 \= n(n-1)/2. In the previous case, all of these are concordant, so this is 1/1. And if it’s discordant, that’s \-1. The sample estimator of Kendall’s τ is $$ \\tau = \\frac{2 \\sum_{i=1}^{n-1} \\sum_{j=i+1}^n \\operatorname{sign}((x_i - x_j)(y_i - y_j))}{n(n-1)} $$ So basically you’re enumerating all these pairs, and then dividing it by the known bottom, and flipping the 2\. What are the dependency limits? * Independence: τ \= 0 * Perfect Positive (Maximum): τ \= 1, or τ \= \-1 Minimum dependence does not exist, it converges to 0\. For a continuous x, y, Kendall’s tau should be $$ \\tau_{xy} = 4\\int_0^1 \\int_0^1 C(u, v) dC(u, v) - 1 $$ ### Spearman’s Rho Linear Pearson correlation for rank correlation $$ \\rho_S = \\frac{\\sum_{i=1}^n (R_i - \\bar{R})(S_i - \\bar{S})}{\\sqrt{\\sum_{i=1}^n (R_i - \\bar{R})^2}\\sqrt{\\sum_{i=1}^n (S_i - \\bar{S})^2}} = \\frac{12 \\sum_{i=1}^n (R_i - \\bar{R}) (S_i - \\bar{S})}{n(n^2 - 1)} $$ expression for min of ranks for x and y, rbar sbar Furthermore, for a continuous x and y. $$ \\rho_S = 12 \\int_0^1 \\int_0^1 C(u, v) du dv - 3 $$ This is important because for each copula it’s going to be dependent, as a function of those parameters, it’s going to have different values in the Pearson, Kendall or Spearman value. There are some advantages to rank correlation, which is that it’s able to capture nonlinear dependence, but it has some disadvantage as well. ## Tail Dependence Rank correlation does not consider the absolute magnitude of the realizations and cannot capture extremely joint behavior. From a rank perspective, it’s increasing or decreasing, or concordant or discordant. For this reason, you can’t look at just one measure of dependence. This refers to the probability of the joint occurrence of these events, of being in the tails. Visual in Notes page 30\. **Upper Tail Dependence Parameter:** $$ \\lambda_u = \\lim_{u \\rightarrow 1} \\mathbb{P}(Y > F_Y^{-1}(u) | x > F_X^{-1}(u)) $$ where f is inverse marginal We’ll say u is the probability of default. Since u is large, This is the conditional probability that y is in the tail, given that x is in the tail. We calculated in the conditional in the survival probability in the previous class, so it’s measured similarly. If λ\_u \> 0, x and y are upper tail dependent. Similarly for lower tail dependence: $$ \\lambda_L = \\lim_{u \\rightarrow 0} \\mathbb{P}(Y < F_Y^{-1}(u) | x < F_X^{-1}(u)) $$ And if λ\_L \> 0, then x and y are lower tail dependent. The Gaussian copula doesn’t have tail dependence. ## Neural Networks and Credit Ratings - URL: https://sharifhsn.dev/blog/computational-methods-week-13/ - Structured data: https://sharifhsn.dev/api/posts/computational-methods-week-13.json - Description: This is an idea that was created in the 50s, exploded in the 90s, where it was called neural networks. A lot of development on the brain happened, so ANN was a term to distinguish … - Date: 2025-04-22 - Exact published timestamp: 2025-04-22 - Topics: Computational Methods, Neural Networks, Backpropagation, Credit Ratings, Volatility Estimation - Categories: Computational Methods - Source: Computational Methods in Quantitative Finance - Source URL: None ## Artificial Neural Networks (ANN) This is an idea that was created in the 50s, exploded in the 90s, where it was called neural networks. A lot of development on the brain happened, so ANN was a term to distinguish from the completely different human neural network. There were a lot of movies about AI, grants given, and they all failed. In 2000, every grant would be rejected because nobody believed it was intelligent. The problem is this stuff is very prone to overfitting. It fits the data you have really well, but trying to predict something else will fail miserably. It cannot move outside that realm. In 2020, they used this to learn the vocabulary. It’s based on the same technology, they put several neural networks together and learned the English language. Turn sout that it works because you can overfit that no problem, because everything you can ask fits in the sample. That’s the principle of generative language. There are some morons like Sam Altman and Elon who don’t understand the technology and think it’s really intelligent. You can only do well in your overfitted universe. You can’t do well outside of that. GPT-5 was supposed to do judgement and reasoning, but it can’t do this on the limits of technology, a retrieval augmented mechanism. Today, we’re not talking about any of that. This course is about the bottom of it all, how exactly do you bild a neural network? ANN is a very general framework, everything that has this kind of mechanics, with lots of designs. LSTM (long short term memory) is an example. ## Multilayer Perceptron The simplest possible ANN is this, which is inspired by brain function. Today we will describe just a single layer, but I’m going to explain how it works completely. At the core, you have $$x \\in \\mathbb{R}^d$$. This is the number. You cannot do any of this with qualitative values. This is why LLMs are built on embeddings, which are numbers. And you also have $$y \\in \\mathbb{R}^k$$. You observe x and y, and you want to know the connection. If I input x, how do I get y? The oldest problem: I have some function y \= f(x), and I want to know what it is. This is the one variable description. With multiple variables, you start talkin about regressions, ANOVA, etc. These are all functions that relate x to y. But all of these things in typical regression are linear, which means f is a linear function. What does linear mean exactly? $$f(x) \= ax$$ In terms of matrices, if $$x \\in \\mathbb{R}^d$$, then $$f(x) \= Ax$$ where A is a matrix $$A \\in M\_{k \\times d}$$ times the x vector $$d \\times 1$$. The whole problem here is to find the A that will give me the output y. This is distinct from an **affine** function. This linear function is constrained so that x \= 0 always gives me 0, (0, 0\) is valid. In general, affine looks like $$f(x) \= Ax \+ b$$ This b makes it affine. Instead of estimating A matrix and b vector, I estimate both. This is regression, nothing fancy here, I’m just making it look complicated, but it’s really nothing. There is a point in me doing this. These are the building blocks for the neural network. The next question in the 50s was this. What is the idea? I am going to somehow measure the distance between my output y and the predicted thing $$\\hat{A}x \+ \\hat{b}$$. This is kind of like a difference for y\_observed \- Ax \+ b, where A and b are observed. I somehow want to minimize this distance. What does that mean? The simplest thing is Euclidean distance. These are all points in $$\\mathbb{R}^k$$, so you can take the sum of squares. You can also take absolute values, but it doesn’t work that well because it’s not differentiable. The minimization happens when you take the derivative equal to 0\. Once I have the minimization, I can obtain the value Ahat and bhat. In statistics we take least squares regression, that’s what this is. Remember that Ax \+ b is f(x). I am literally taking the difference between y and f and finding out f. But what if the function f is nonlinear? Then I’m screwed. The least squares is very simple to derive because the other term will disappear when you take the derivative of a particular term. The advantage in the 80s was to come up with this function. ## The Idea Here are the things we have to work with: $$x\_1^1 \\ldots x\_d^1, y\_1^1\\ldots y\_k^1$$ Typical the models we use are output one dimensional, but nothing prevents you from making it k-dimensional And then all the way up to $$x\_1^n \\ldots x\_d^n, y\_1^n \\ldots y\_k^n$$ I have n observations. I have 40 people in the class and I start measuring their x values, like height. Then I monitor something internal to predict, their cholesterol level. There was a commercial which was the IBM intelligent machines. In Manhattan, for this coffee place, they discovered as they train people to order more puffs, then the IBM machine looked into the data, and they sold more puffs. I hated that commercial because it implied that the machine somehow went int o the data nad saw that connection. You are literally looking at all the k products. First of all, it has to be logical. Machines cannot discover connections. You the researcher will hypothesize the connection, and the machine will check it for you. The question is, how do we relate x to y? This is howt he neural network works. (Full diagram in notes) Let me take the components x1… xd This is for a generic input. We talk about random variables in financial engineering. These are observations. We are making a relationship between the random variables in the input and the output y1 … yk. How do you make this connection? We are going to form this hidden layer. In the simplest case, just one layer. We will call it h1… he. I’m going to make a nonlinear activation function. I’ll combine all of the xs into a nonlinear relationship with the hs. Then I’ll combine these hs in a nonlinear relationship with the ys. How do you do this? First you take for every h\_i, you consider the function g\_1. The difference between linear and affine, we will construct an affine relationship. I will add an x\_0 \= 1 here. This is a free term, this constant. I will add some weights which connect these xs to the hidden layers. w\_01^1, w\_11^1, … w\_d1^1. This h\_i will be a function of $$h\_i \= g\_1(\\sum\_{j=0}^d w\_{ji}^1 x\_j)$$ This relationship is really not important. I am literally taking a linear relationship, this function g\_1, and I’m applying the affine to it $$= g\_1(w\_{0j} \+ w\_{ij}x\_1 \+ \\ldots \+ w\_{dj} x\_d$$ I’m taking all of the inputs times this weight. And that’s just for one node. I need to do it for all layers because I will combine them when I reach y. The y will be another function g\_2. We will add another h\_0 \= 1 at the top, a constant. So that this function will be $$y\_k \= g\_2(w\_{0k}^2 \+ w\_{1k}^2 h\_1 \+ \\ldots \+ w\_{ek}^2 h\_e$$ Literally, you are taking a linear combination of the weights, times the nodes, and you put it into the activation function, and put it into the nex tithing. At the end, you have these output y from a mountain of inputs. So what do we do next? ## Activation Functions We mentioned g\_1 and g\_2. What are those? These are called **activation functions**. If you’ve taken statistics, or FA590, at the end of the class, you do something called logistic regression. You are associating real numbers with probability. Because you need to map them continuous to continuous, since you can’t really map from continuous to discrete. Instead of mapping x into the outcome, I will map x into something which maps into \[0, 1\], this reduced region. The typical activation functions are the following. Technically you can use anything, but you will usually depend on one of seven fundamental functions, something every math student learns. Exponential, tirgonometric, polynomial, etc. (Graphs in notes) Examples include: - hyperbolic tangent: $$g(x) \= \\tanh(x) \= \\frac{e^x \- e^{-x}}{e^x \+ e^{-x}}$$ This maps into \[-1, 1\] If x is multidimensional, then replace this with something multidimensional You can’t use cosine because it will explode at some values. This is what these things are trying to do. If you have x in the left region, it maps into a negative value, and the right is positive. So you can literally by playing with the weights, you can make your resulting point, which is a linear combination of inputs times weights, you can make it be either left or right. You can guide your output to y to be more positive or more negative. - logistic: $$g(x) \= \\frac{1}{1+e^{-x}}$$ It’s kinda similar to tanh, but the difference is that it goes into \[0, 1\]. The reason it’s called logistic and it’s used in logistic regression, because you’re mapping real numbers into probabilities, and this is the correct domain for probabilities. - Rectified Linear Unit **ReLU**: $$g(x) \= \\max (x, 0\) \= x\_+$$ This is used in machine learning quite extensively. My mom decided to call me Ionut, so I’m only known by that one name. But the computer science decided to call this The ReLU gate is used to delete negative numbers. It’s used a lot in computer engineering. - Softplus: $$g(x) \= \\log(1 \+ e^x)$$ ReLU is called xplus. But it has a problem at 0, where it’s not derivable. So this makes it smooth. It’s nonzero everywhere and derivable, but it has the same general form. This is the activated function, you can use it in both places, construct it however you like. It’s really important that you don’t do it with trial and error. People fire up pytorch and use defaults, and think it’s good. Obviously each of these activation functions have their own meanings. You need to know when to use one or the other. Typically, the probability stuff uses logistic for the y. Typically, you want to predict some kind of number. Sometimes SoftReLU is used, sometimes regular ReLU, I don’t know which one is better. Depends on the output’s relationship with the input. ## The Weights How many weights, and how do estimate them? How many weights is pretty simple. I have d input, and k output. In that case, it depends on this internal layer. We’ll say the hidden layer has e nodes. Then I have d \+ 1 inputs (adding the affine part). Each of those d \+ 1, I weight each of them to connect to each of the e nodes. That’s (d \+ 1)e. That’s just for w\_1. Now I have to do the same thing to connect with k output ys. e \+ 1 affine term, and k of it. w\_2 has (e \+ 1)k terms. So it’s total (e \+ 1)k We’ll say x is 20 dimensional, we have 20 parameters, and y is one dimensional, with one output. What is the typical number of nodes in a hidden layer? 64 ‘ That would make this calculation 1409 weights. That’s an immense number of weights. In regression, if we have 20 inputs, then we have 22 parameters. If this is a class, and I add class characteristics, and I have years, and I have only 1200 cases, and I fit a neural network, what’s going to happen? I will get a perfect fit. Not only that, it will not be unique, because there are too many weights compared to observations. What about 10k cases? It will still be overfit. You will need millions of observations. That’s why when you’re doing discriminants between images, you need thousands. There are advanced methods to deal with this. If you want to create a model of a person (this is an old thing, the computer vision problem). When I ame to Stevens, I did CV research for a while, in it searly days. When you plugged in 20 algorithms, 19 would not even work, just give wrong answers. They were based on this overfitting data that you have. Today is 20 years later, very mature CV, things that appear and disappear, all these techniques. You can get the shape from shading problem by lighting it differently, and get a model of a person. A friend of mine has this problem 2D to 3D. All of this is set in this particular domain. ## Backpropagation This is complicated, I’m not going to explain in detail. But I will explain the fundamentals. We have these weights that we’re trying to estimate. I know if I put my weights with my function, I get a candidate. So we get something called a Loss Function. We already have one here, which is the least squares. There are two cases, if y is discrete or continuous. For the height of a student it’s continuous. If you output a gender, it’s discrete. Squared error loss is good if it’s continuous: $$\\sum\_{i=1}^n (y\_i \- \\hat{y\_i})^2$$ yhat\_i is the output of the neural network. $$\\sum\_1^ |y\_i \- g\_2|\\sum w\_i^1 g\_1(\\sum w^2 x))^2$$ This function contains all 1000 observations, and minimizes it. The backpropagation does not minimize all the 1000\. It minimizes w2 first, then it minimizes w1 with respect to that, propagating the weights backwards, and then doing it again. This is the simple part. The complicated part… I learned in my other class that I should use mean squared error. Who gives a shit. Some of you may say that some are more important than others. You can definitely do that. You might say it’s very important to fit particular ys, and make it a weighted sum. But most of the time, it basically looks like this, with some kind of Euclidean distance. What do you do in the discrete case? They use something called Cross-Entropy which is bullshit. The loss function for the discrete variables is called this, and it’s hard to explain. So first I will explain something simpler. If we have discrete, we are going to output either male or female. We’re going to create two nodes. We will have probability of one, and 1 \- probability of the other. If you have multiple outputs, you use a hot-cold encoder, which maps this into a multi-dimensional space. The upshot is that you have to compare the observed discrete distribution and what you output, which is going to be probability. The function mapping is of numbers. In any of these ML techniques, you will take those numbers and calculate the probability distribution out of them. Then how do you compare this probability distribution from the model, with the true probability distribution. Let’s say I have three outputs, green yellow red. If my object that I have is a combination of them, I will write down the combination. That’s the distribution I’m looking for. Most of the time, the categorial thing I’m outputting is one of them. 30s, 20s, teens. When I’m looking at one picture, the actual observed will be 0 0 1 0 0\. The probability distribution is 1 that belongs to this category and 0 to not others. So we nee dto calculate the difference between the output and this distribution. We will use cross entropy, but it’s hard to understand so we’re not using this. “If p and q are two probability distributions, and they have the same support. The rv is characterized by probability and outcome. The support is the domain of outcomes.” For continuous, this is not useful, because we have MSE. But it can be useful for discrete. $$H(p q) \= \-\\mathbb{E}^p(\\log q) \= \-\\sum\_{i=1}^k (\\log q\_i)p\_i$$ What does that mean? It’s outcome times probability. It’s kind of interesting. They have to be the same outcome to calculate this. So how does this work in terms of ML techniques. I have three outcomes: red, green, blue. One of the distributions is going to be 1 0 0, the other one from the neural network will be a b c, which corresponds to red green blue. So we have to pair these. That’s the cross-entropy. Instead we will look at **Kullback-Leibler Divergence**. This is defined as $$\\mathbb{E}^P(\\log\\frac{P}{Q}) \= \\sum\_{i=1}^k \\log \\frac{p\_i}{q\_i} p\_i \= \\sum p\_i \\log p\_i \- \\sum p\_i \\log q\_i$$ IT’s pretty much the same thing as entropy in practice. The first term is only in p, which is typically a constant, it’s what I know. This is simple to understand. If p\_i is close to q\_i, this is close to 1\. The logarithm of 1 is 0\. So it’s basically it’s a bunch of small parts of 0.This looks likeaa distance between p and q. But it’s a divergence, not a distance. Distance has three properties. distance should be commutative, reflexive, and additive (triangle inequality). This only follows the first two. If something is equal to 0, then the two are the same. But, if something is tiny different from 0, and in other experiment you have the same tiny number, you cannot say that this cross function is the same. The distance itself is meaningless. Distance, I can measure 300 km here, 300 km in Romania, which is the same. Divergence, hell no. ou measure something here, the Trump tariff will be different than in Romania, because that’s not a distance. This is neglected in all of CS literature. It’s a lot of wishy-washy here. These numbers are special numbers, they’re between 0 and 1, and all sum up to 1\. That’s why you can’t use MSE. If one is distant, the other two are close. In the most common situation, the output is one and a bunch of 0s. It turns out that thing that maximizes it will give you 1 for each of the outcomes. That would be completely overfitting. READ ABOUT LOGISTIC REGRESSION. The way that it’s done is the same as it’s done with this, except you’re using logistic vs Kullback-Leibler or Cross Entropy. ## Logistic Regression q\_i is the output of th efunction you get from the node. $$q\_i \= \\frac{1}{q \+ e^{-w^ix}}$$ If you’re going directly from x to y without a hidden layer in between. There are weights that correspond exactly to this particular outcome. Typically you take q\_i as $$q\_i \= g\_1(q\_1, \\ldots)$$ For the output, you want to make sure bigger is bigger and lower is lower. ## A case study in corporate credit rating Basically the idea is, we have a credit rating. A company raises money by issuing stock, or debt. A debt is a loan, and repayment of the loan depends on credit rating. If the rating is high, you can get a low repayment rate. Who does the rating? S\&P, Fitch, and Moody’s. But there are many more smaller ones, DB Morningstar, that hire a lot of our students. In India, these companies, do not operate there. You have consultant companies which work with S\&P to help with their models, and move them to India to use them there. Well, at least they do this in Europe. The bank will hire a consultant that implements a Moody’s model and gives you the answer, it’s very bad. In practice, you have all the ratings. We looked at 16 years, some companies started more recently, some later. 62 companies, with 297 financial features, 297 xs. One random variable is the output, but it’s all 24, because you want to create an output for each of the likely outcomes. This is basically talking about the various techniques we used. How do we evaluate algorithms? How good is it? Once you train you get the weights. So you have to do cross validation to test it. Typically we have two data sets, training and testing. You typically have three regions in finance. In CS, you split into training and testing. Testing you just use for output, that’s it. Testing you split into this K-fold cross validation techniques. You’re always dealing with hyperparameters, should I have one hidden layer or multiple, 64 nodes, or 128, or 50, how many? The way to determine this is to split the data into 10 pieces, use 9 pieces to construct your weights, and then use that to predict the 10th part, and then see how good it is. Then you pick your hyperparameters which you can use. Estimation of actual parameters is done with the other side of the data. Now you have your weights, you can plug in your features, and get a credit rating. You use these metrics, which have nothing to do with cross entropy baloney which is used to calculate weights. Once you have the weights, you’re concerned with accuracy. I have my company which is rated A, which has an output number. Is it the same as the observed number? You can do accuracy for the entire dataset. I have 100 points, and 50 of them I guessed the credit rating exactly, the other 50 I did wrong. So my accuracy is 50%. We also have something called recall. Of those 100 companies, 20 were AA. Of those, the model only predicted 15 of them. 5 of them were predicted as being something else. That’s recall. Recall can only be done per category, and accuracy is done overall. Precision is the same as accuracy, but only for those categories. You have 100 observations, accuracy for 100 points. My algorithm predicted 25 AA. There were 20 AA. Of those 20, 15 were actually AA. My precision is 15/25, my recall is 15/20. Math is in corporate credit rating slides for Precision and Recall. You have two numbers, how many of them were guessed correctly, and how many of the guesses were correct. Which one is better to describe this? The F1 score is the harmonic mean of precision and recall. There are 3 means, arithmetic, geometric, and harmonic. Arithmetic is largest, and harmonic is smallest, always. This is as conservative as possible, so it picks harmonic, the smallest possible average of the two. This is for only one category. You can do **macro-averaging**, regular averaging, or **micro-averaging**, which is the same as accuracy. There’s not that much complication here. You just need to understand the fundamentals. The skill for interviews: you need to explain the model. GIVE EXAMPLES. ## Shanshan’s Volatility Estimation Goals: show that ANN can model complex nonlinear relationship, like a function with no closed-form solution And how to do this with PyTorch Training and testing process, will be completely We set up with 5 input features, one output . Hyperparameters we set up ourselves. There are two hidden layers with 64 nodes in each layer, ReLU is used. Hidden layer can choose different activation functions. ReLU is common. For output layer, I used ReLU because Volatility is always positive, and can be \>1. The W1 and b1 parameters are the initialized parameters, which will be updated during training. Defining PyTorch neural network is done by extending nn.Module in a class. You define a function Sequential, with a series of nn.Linear and the activation (nn.ReLu). The number of input and output features for each layer is given as parameters to Linear each layer. Then you run the sequential for a given input. Dataset prep: you do this beforehand to understand the data. training data and test data are split to evaluate the model convert the dataset into tensor, wrap in Dataset class, allows you to use DataLoader. You can do batches for efficiency, and shuffle training data for automatic cross validation. Use Adam optimizer to automatically adjust learning rates, faster than gradient descent. Compute training MSE loss to monitor progress Evaluation MSE loss for overfitting Testing MSE loss for results How to prevent overfitting? When it learns the noise and random fluctuations instead of general pattern. Use an evaluation set during training. If training loss and evaluation loss diverge too much, that’s a problem. Ways to deal: * Early stopping. Orange line stays, but blue line is decreasing. Then stop the epochs as the point when the blue line passes the orange line. * Regularization. Penalize large weights. If the model is very complex, it has the space to learn the noise and random fluctuations rather than general pattern. Adam has (, weight\_decay=1e-5) ## XVA and Counterparty Risk - URL: https://sharifhsn.dev/blog/xva-and-counterparty-risk/ - Structured data: https://sharifhsn.dev/api/posts/xva-and-counterparty-risk.json - Description: The XVA section adds counterparty and funding effects to an otherwise clean derivative value. CVA is the expected loss from counterparty default on positive exposure; FVA captures … - Date: 2025-04-21 - Exact published timestamp: 2025-04-21 - Topics: FX, XVA, Counterparty Risk - Categories: FX - Source: FE-635 \| Risk Engineering - Source URL: None The XVA section adds counterparty and funding effects to an otherwise clean derivative value. CVA is the expected loss from counterparty default on positive exposure; FVA captures the funding cost of carrying an uncollateralized position. The exact decomposition depends on collateral, netting, and the institution's convention. The calculation is path dependent: simulate or approximate future exposure, combine it with default probabilities and recovery, and discount the expected loss. Netting sets and collateral agreements can change the exposure more than a small shift in a market input. The class notes use XVA to connect pricing and risk governance. A model output is meaningful only when the legal agreement and the exposure definition are the same ones used by the desk. ## Why Federal Reserve Independence Matters - URL: https://sharifhsn.dev/blog/linkedin-2025-04-18-why-federal-reserve-independence-matters/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-04-18-why-federal-reserve-independence-matters.json - Description: Trump is threatening the independence of the Federal Reserve. But why does this even matter? And what does the Fed even do? - Date: 2025-04-18 - Exact published timestamp: 2025-04-18T16:23:13.160Z - Topics: Central Banking, Monetary Policy - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7319032768905064449/ Trump is threatening the independence of the Federal Reserve. But why does this even matter? And what does the Fed even do? The role of central banks in the economy is a privileged one. They have the ability to manipulate the supply of money, and more importantly, they set the overnight interest rate that banks will lend money to each other at, which cascades into the interest rate for borrowing money across the whole economy. The Federal Reserve has a dual mandate, which Chairman Powell stresses in every meeting: to meet target inflation and unemployment. Generally, lowering interest rates lowers unemployment. When businesses can borrow money more cheaply, they’re more willing to expand and hire. However, this also increases inflation. Demand for labor comes with a demand for everything else, and if supply can’t meet it, prices will rise. This was seen most dramatically when the Fed dropped interest rates to zero after the COVID pandemic. It prevented mass layoffs, but the increased demand from consumers combined with a lack of supply from pandemic-era supply chains led to record inflation. The Fed’s decisions reverberate through the market, and a lot of money rides on predicting this. In fact, you can see exactly what the market thinks the future interest rates will be by the Fed Funds Futures curve. CME’s FedWatch tool provides a nice visualization of this in probabilistic terms. As you can see in this graph, the market currently thinks that the Fed will continue to pause interest rate changes at the next meeting in three weeks. This is because the Fed is conservative. The organization is composed of experts who deeply understand the impacts of their decision on the economy, and will not threaten it for short-term economic reasons, especially in times of instability. The same cannot be said for politicians. They are motivated to keep their constituents happy in time for the next election. This biases them in favor of lowering interest rates, which will heat up the economy in the short term, even if it might cause inflation in the long term. If the Fed has to bow to political pressure, the probabilities you see would shift dramatically to the left, since Trump has been loudly proclaiming that he wants to see interest rates cut. Such a move would be completely unprecedented and would undermine trust in the dollar as a whole. The falling stock market, rising treasury yields, and falling dollar all imply that institutions are already seeking to exit US markets, and this would only exacerbate this process. Do you think Trump will fire Powell? Tell me in the comments. Read more at Bloomberg: [https://lnkd.in/gBQviQzy](https://lnkd.in/gBQviQzy) ## Dependence and Copulas - URL: https://sharifhsn.dev/blog/advanced-derivatives-week-12/ - Structured data: https://sharifhsn.dev/api/posts/advanced-derivatives-week-12.json - Description: Relevant variables for these credits - Date: 2025-04-17 - Exact published timestamp: 2025-04-17 - Topics: Credit, Dependence, Copulas - Categories: Credit - Source: Advanced Derivatives - Source URL: None ## Dependence cases Relevant variables for these credits **Page 4** Let’s look at, let’s say, minimum dependence: In what situation might we have such minimum dependence? The k where we have the lowest probability that we have both names defaulting. Both of these credits have an idiosyncratic component and a market exposure z. So one way we can achieve minimum dependence is to have an opposite exposure to z. If we have a large negative exposure to z, we’ll say beta is large. For credit j, we can change the exposure to market variable z, and instead of having a large negative number, it becomes a large positive numbers, and we lower the probability of default. For minimum dependence, if we take it in the **limit**, for β\_i \= \-β\_j In the limit, we can take β\_i \= 1, β\_j \= \-1 One has 100% exposure to z, the other has negative exposure to z, and both of them have zero idiosyncratic component. So correlation is inverse, ρ \= \-1. In this case, such probability is dependent on the survival probability of the names, \((1 - Q_i(T) - Q_j(T))_+\) The joint probability of default is zero as long as Q\_i(T) \+ Q\_j(T) \> 1 Independence means correlation is 0\. And to do that, we have to remove market exposure. So in the limit, we take β\_i \= β\_j \= 0, which means the correlation is zero, it’s entirely idiosyncratic. Maximum Dependence is when β\_i \= β\_j \= 1, maximum market exposure. It’s the minimum of either 1 \- Q\_i(T), 1 \- Q\_j(T). If you want to illustrate this, **Slide 5** ## Euler–Milstein and Monte Carlo Extensions - URL: https://sharifhsn.dev/blog/computational-methods-week-12/ - Structured data: https://sharifhsn.dev/api/posts/computational-methods-week-12.json - Description: We will expand on the variance reduction techniques, we will repeat a couple things from last week. - Date: 2025-04-15 - Exact published timestamp: 2025-04-15 - Topics: Computational Methods, Euler–Milstein, Monte Carlo, Control Variates, Correlated Processes - Categories: Computational Methods - Source: Computational Methods in Quantitative Finance - Source URL: None ## Euler Milstein We will expand on the variance reduction techniques, we will repeat a couple things from last week. But before we do that, we should talk about **Euler Milstein** or Euler Maruyama. \[People keep talking. It hurts Florescu because he has a disease called ADHD.\] This is a better approximation. I’m going to show you this because it involves applying Itô’s formula a bunch. Here is the stochastic process we can approximate. Euler works for any stochastic process, but Milstein has to be homogeneous, like so: $$dX\_t \= \\alpha(X\_t) dt \+ \\beta(X\_t) dW\_t$$ These coefficients are not functions of time, so they are homogenous. The iea is to use Itô for bothα and β. They are functions of stochastic process X\_t. THerefore I can use Itô for both. I get $$d\\alpha(X\_t) \= \\alpha’ (X\_t) dX\_t \+ \\frac{1}{2} \\alpha’’ (X\_t)(dX\_t)^2$$ This is Itô, so we are lacking the dt, that term would screw up our calculations. Now we substitute X\_t here, so we get The dX^2 gets eliminated $$= \\alpha’(X\_t) \\alpha(X\_t) dt \+ \\alpha’(X\_t) \\beta(X\_t) dW\_t \+ \\frac{1}{2} \\alpha’’(X\_t) \\beta^2(X\_t) dt$$ If I just group these terms and drop X\_t all over the place to make it easier to write: $$=(\\alpha’ \\alpha \+ \\frac{1}{2}\\alpha’’ \\beta^2)dt \+ \\alpha’ \\beta dW\_t$$ Now, if we do the same calculation/derivation for the β, then we basically get something very similar. We’re not going to go through the entire calculation. You can derive this yourself, $$d\\beta(X\_t) \= (\\ldots) dt \+ \\ldots \+ dW\_t$$ Writing in integral from, Now we substitute α and β from the same formula. So there are two integrals: ~~~text X\_{t+\\Delta t} \- X\_t \= \\int\_t^{t+\\Delta t} ~~~ When we substitute, we will get four terms, two terms for α and β each. We will get terms such as $$dsdu \\sim O(\\Delta t^2)$$ We will change the letters to make sure they’re right, based on what our dummy variables are $$dsdW\_u \\approx dudW\_s \\sim O(\\Delta t^{\\tfrac{3}{2}})$$ Because ds is order Δt, and dW\_t is order √Δt Then there is $$dW\_u dW\_s \\sim O(\\Delta t)$$ FOr the reason just discussed. Once I substitute everything, I will neglect the first two orders, and will be left with terms with just du and ds $$du$$ $$ds$$ The existing equation then simplifies a lot $$X\_{t+\\Delta t} \= X\_t \+ \\alpha (X\_t) \\Delta t \+ \\beta (X\_t) \\Delta W\_t \+ \\int\_t^{t+\\Delta t}\\int\_t^u \\beta\_s’ \\beta\_s dW\_s dW\_u $$ And there’s another integral term you can see here. You need to take the increments of Brownian motion and such, but then you can show that the integral term is equal to the $$= \\beta\_t’ \\beta\_t \\frac{1}{2}(\\Delta W\_t \- \\Delta t)$$ This gives the Euler Milstein scheme. The X\_t+Δt is what you’re approximating. The first part is the regular Euler, and then the integrals are the Milstein part. If the model $$dX\_t \= \\alpha(X\_t) dt \+ \\beta(X\_t) dW\_t$$, then the Euler Milstein scheme is $$X\_{t+\\Delta t} \= X\_t \+ \\alpha(X\_t) \\Delta t+ \\beta(X\_t) \\Delta W\_t \+ \\frac{1}{2}\\beta’(X\_t) \\beta(X\_t) (\\Delta W\_t^2 \- \\Delta t)$$ And by W you introduce a normal variable multiplied by $$Z \\sim N(0, 1)$$. The ΔW is created by $$Z\\sqrt{\\Delta t}$$. The other one is $$(Z^2 \- 1)\\Delta t$$ when the thing factors. For the same increment, you put it in two places, not just one places. However, there is a possible issue, which might be the derivative. If the β function, the volatility part, if it’s complicated, how do you calculate the derivative? That might be hard. There’s another way to deal with this, a scheme called Runge-Kutta, a generalization where you calculate the euler part, and plug it into the beta, then you calculate the finite difference as an approximation of the derivative. One more thing to mention. Which I shouldn’t, because it’s from the homework. Let’s have an example. We have a process $$dY\_t \= \\kappa(\\bar{y} \- Y\_t) dt \+ \\sigma \\sqrt{y\_t} dW\_t$$ This is the CIR process. In this process, I have my $$\\alpha(x) \= \\kappa(\\bar{Y} \- x)$$ Then you have beta $$\\beta(x) = \\sigma \\sqrt{x}$$ If we now substitute in this formula, we have to calculate the derivative. $$\\beta’(x) \= \\frac{\\sigma}{2\\sqrt{x}}$$ Now if we do Euler Milstein: $$Y\_{t+\\Delta t} \= Y\_t \+ \\kappa(\\bar{Y} \- Y\_t) \\Delta t \+ \\sigma \\sqrt{Y\_t} \\Delta W\_t \+ \\text{ milstein correction: } \\sigma \\sqrt{Y\_t} \\frac{\\sigma}{2\\sqrt{Y\_t}} (\\Delta W\_t^2 \- \\Delta t)$$ Here, the two sqrtY\_t cancels, so it becomes sig^2/2. The other thing to show is that if you look for example, this process. $$dX\_ \= e^{X\_t} \\cos X\_t dt \+ 0.7dW\_t$$ Here, you have this constant $$\\beta(X\_t) \= 0.7$$ So $$\\beta’(X\_t) \= 0$$ Therefore there is no Euler-Milstein correction, because it relies on multiplication by the derivative. Like for example in GBM. ## Variance Reduction Redux More of the idea to use on HW 4\. One additional note will be given on antithetic variate. We discussed the CLT and the basis of the whole thing, and the variance being smaller gives you better estimates. However, this is more complicated than that. **The better method is not always the one with smaller variability of the sample paths.** We also need to consider the time to generate paths. If you have a Monte Carlo technique with less variability, but it takes one minute to generate a path, then it could be that another method with much higher variability that can generate 1 per second, is better, just by raw brute force. The paper mentioned is really good, the basis of the Monte Carlo book, which expands on the paper. Boyle is a Canadian professor from University of Waterloo, in 1997 they met, he did a summer school. He was drunk all the time in the morning lectures. Glasserman is a friend of the show as well. Third guy Brodie sucks. If method 1 has variance σ\_1^2, and b\_1, they have a term called “work”, which could be time, but could be other stuff, b\_1 is the work to generate one replication of the final parameter, the one you’re trying to estimate, then we need to look at $$\\sigma\_1^2 b\_1 \< \\sigma\_2^2 b\_2$$ He has a better expression, but I’m showing the simple thing. Why multiplication? There’s a reason If you look at this perspective, and if you rewrite this as $$\\frac{\\sigma\_1^2}{\\sigma\_2^2} \< \\frac{b\_2}{b\_1}$$ That’s when you would prefer one method over another. **Antithetic note**: Reminder: The way it works is you create one normal for one path, and use the negative of that normal for another path. For antithetic variates, one replicate is $$\\frac{c\_i \+ c\_i^a}{2} \= \\bar{C\_i}$$ The actual estimate is the average of the 2\. This is not necessarily important for the final pricing, because if you take the general average, it doesn’t matter, but it’s important for the variance estimate, for confidence intervals. This is because these paths are not independent, they’re very related. So you would use something like $$\\text{Variance} \= \\frac{1}{n-1} \\sum\_{i-1}^n (\\bar{C\_i} \- \\bar{C})^2$$ This is important because each sequence of n is independent. You might think I should iterate over 2n elements instead of n. and do both the regular and antithetic. But this one will show less than it really is. ## Control Variates for the Asian Option Glasserman’s explanation is better than mine, I will follow him. The one with delta hedging and Asian option. There’s nothing wrong with delta hedging or Asian option, it’s just that it needs to be explained where it’s coming from. The control variate for Monte Carlo, what it does, it should be called “use what you know”. You know, for example, that the option that delta hedges, you’re using the concept that if you do this very fast, the two values should be the same. Similarly, for the Asian option, we’re using what we know. $$P\_A$$ is the price of an Asian option based on Arithmetic average, which is what is encountered in practice. $$P\_G$$ is geometric. This P\_G has a formula, and for any variation on Asian options it’s very easy to get a geometric formula. We’ll say we have a formula, if you give me characteristics and μ and σ, you get an exact number. With this, let $$\\hat{P\_A}$$ be the value calculated using a single path. It doesn’t matter what you use, just one single path, where you calculate the value of the Asian option based on this average of all these paths. And we’ll do the same for geometric $$\\hat{P\_G}$$ We don’t need this for geometric because we have the formula. **We are using the same path for both of these**. I do know that $$\\mathbb{E}\[\\hat{P\_A}\] \= P\_A$$ On expectation, I get the true vlau eof my option. I also know $$\\mathbb{E}\[\\hat{P\_G}\] \= P\_G$$ If we subtract, we get $$P\_A \- P\_G \= \\mathbb{E}\[\\hat{P\_A} \- \\hat{P\_G}\]$$ This gives you a very natural kind of estimate. $$P\_A \= P\_G \+ \\mathbb{E}\[\\hat{P\_A} \- \\hat{P\_G}\]$$ We can create a Monte Carlo path using this control variate: $$\\hat{P\_A}^{cv} \= \\hat{P\_A} \- \\hat{P\_G} \+ P\_G$$ If I create a new path and average this, then it should give me on average this difference You can write it like this $$= \\hat{P\_A} \+ (P\_G \- \\hat{P\_G})$$ That parentheses statement is the control variate. I have my original path with Euler Milstein, then I control it. For every path, I obtain the difference between the true value and the value of geometric from that particular path. But is this better than just using $$\\hat{P\_A}$$? If we want, we can calculate the variance. So the variance should be less. And remember that $$P\_G$$ is just a number, a constant with no variance. $$\\mathbb{V}\[\\hat{P\_A}^{cv}\] \= \\mathbb{V}\[\\hat{P\_A}\] \+ \\mathbb{V}\[\\hat{P\_G}\] \- 2 \\text{Cov}(\\hat{P\_A}, \\hat{P\_G})$$ Everything after the first term should be negative to give me a better variance. The covariance should be greater than the variance. It’s only worth it if the covariance is large. This brings the next idea. This is the original term plus this term. I can control the size of the difference, which puts a β on the coefficient. That β allows me to make the thing smaller. This is all specific to the Asian option, where there is this arithmetic and geometric thing. But there is no assumption about the stochastic model. β means I’m going to parameterize this: $$\\hat{P\_A}^\\beta \= \\hat{P\_A} \+ \\beta(P\_G \- \\hat{P\_G})$$ I can play around with β in such a way that the resulting variance is the smallest. What is that? We can do the same exact calculation, and minimize the result with respect to β. $$\\mathbb{V}\[\\hat{P\_A}^\\beta\] \= \\mathbb{V}\[\\hat{P\_A}\] \+ \\beta^2 \\mathbb{V}\[\\hat{P\_G}\] \- 2 \\beta \\text{Cov}(\\hat{P\_A}, \\hat{P\_G})$$ This is a quadratic expression. It’s a parabola, so the smallest value is in the vertex. The position of the vertex is obtain for $$\\beta\_{\\text{min}} \= \-\\frac{b}{2a}$$ Where the varG is a, and the cov stuff is b. $$= \-\\frac{-2\\text{Cov}(\\hat{P\_A}, \\hat{P\_G})}{2\\mathbb{V}\[\\hat{P\_G}\]}$$ This cancels to: $$= \\frac{\\text{Cov}(\\hat{P\_A}, \\hat{P\_G})}{\\mathbb{V}\[\\hat{P\_G}\]}$$ If you have two variables, the regression is the covariance divided by the variance, so this is the formula for market beta/regression. So if we regress $$\\hat{P\_A} \= \\alpha \+ \\beta \\hat{P\_G} \+ \\epsilon$$, the β is the β. Last week, we learned this was true, but now we know how to get it. **This control variate thing is basically** $$\\hat{P\_A}^\\beta \= \\hat{P\_A} \+ \\hat{\\beta} (P\_G \- \\hat{P\_G})$$ Now we have another problem. This is kinda screwed up. Because you’re using the same paths to estimate β, and the same path to estimate the value of the option. That introduces a bias, and this is calculated in the Glasserman paper. Typically you have n paths, and you set n\_1 paths out to do regression. Then you use n \- n\_1 paths for calculation. The advantage here is to do a proper regression, you don’t need a lot of observations, 100 would be plenty. But for Monte Carlo, you need hundreds of thousands. Hopefully this is more clear, and what I hope you get is that he used the particular relationship that exists in the Asian option, to reason through the whole thing. I’m going to make another expansion to this. We can introduce more control\!\! It’s not really necessary because this existing control variate already gives a good estimate, but this shows you can introduce as many control as you like\! Under risk-neutral equivalent martingale measure, we have $$S\_0 \= \\mathbb{E}^Q\[S\_T e^{-rT}\]$$ If you take the stock price as a martingale, and discount it back, you should get S\_0. Nothing new. This brings up another way to control. We use the following, with β\_1 being our original control variate. $$\\hat{P\_A}^{cv} \= \\hat{P\_A} \+ \\beta\_1(P\_G \- \\hat{P\_G}) \+ \\beta\_2 (S\_0 \- \\hat{S\_T} e^{-rT})$$ Now you have the path, you know what S\_T is, and S\_0 is a constant. You have the same kind of regression of $$\\hat{P\_A} \= \\alpha \+ \\beta\_1 \\hat{P\_G} \+ \\beta\_2 \\hat{S\_T} \+ \\epsilon$$ The constants don’t matter because you’re doing a regression, it just changes the y-intercept, doesn’t impact the β. Technically, it’s σ√t, which is the confidence interval size, the diffusion size. It’s an estimate, work could refer to other things. ## Moment Matching Method This is simple to use, so I’ll mention it, even though it’s useless. If we have $$Z\_1, \\ldots Z\_n \\sim N(0, 1)$$ They should have theoretical mean 0, but the sample mean is not 0\. The idea is to modify the sample to match the theoretical moments. The reason I actually have never taught this method is because it’s kinda stupid, as a statistician. Also, from the raw power of this method, it doesn’t do better than straight Monte Carlo. It does better when you pair it with control variate, where the power comes from the control variate. In the example here, the random variables don’t have mean 0, so instead use $$Z\_1 \- \\bar{Z}, Z\_2 \- \\bar{Z}, \\ldots$$ These now all have mean 0, but now they’re correlated. So it creates problems when estimating stdev, it becomes bad, very tricky for estimating confidence intervals. In the example, let’s say I’m going to price a European option based on GBM. When you do GBM, you don’t have to do all the intermediate steps, you can do all in one step. Because the terminal value $$\\tilde{S\_T}(i) \= S\_0 e^{(r-\\tfrac{\\sigma^2}{2})T \+ \\sigma \\sqrt{T} \\tilde{Z}\_i}$$ After you modify the normal variable with the thing I said. This is fine for European options, but it won’t work for path-dependent options. Confidence intervals are hard to obtain. This is the first order moment matching. You can also do second order moment matching. Say we want to create $$N(\\mu\_Z, \\sigma\_Z^2)$$ I’m trying to create numbers that are normal with this particular target. The usual thing to do here is $$Z\_i \\sim N(0, 1\) \\rightarrow \\sigma\_Z Z\_i \+ \\u Z$$ Multiply to create the desired distribution. But the numbers in the sample will have their own sample mean and stdev. So you have to modify this as $$\\tilde{Z\_i} \= \\frac{\\sigma\_Z}{S\_Z}(Z\_i \- \\bar{Z}) \+ \\mu\_Z$$ Each number is modified by these two numbers S\_Z and Zbar, where S\_Z is the sample stdev. $$S\_Z \= \\sqrt{\\frac{1}{n-1}\\sum (Z\_i \- \\bar{Z})^2}$$ For the random variables in the sample, they will have the desired distribution. Like I said, this is a method from the 90s. I never liked it because it’s slower. All these modifications… And in order to estimate the sample mu and stdev, I have to do all of the simulations first, then calculate the samples, then plug them back, so it’s SLOWER than doing it all at once. And generally, from my experience, improvement is marginal, it’s not particularly useful. Is the sample calculated in one path or all paths? It depends. In this example, you need all the paths. But you can do it for one path if you like. Computationally, when you have numbers that you’re observing, if I have to estimate an average or variance of a certain number of them, theoretically, I can do this. The idea of the GPU is that you have a lot of matrices and you do these calculations really fast. There is a memory of GPU and a memory of CPU. You want to keep everything in GPU memory as much as possible. The problem is that this method aggregates the paths, and does it again, which is impossible with a GPU. ## Additional Stuff Stratified sampling is not the best explained here. It is a statistics technique, check any statistics book they will explain better there. Importance sampling, same deal. I have a chapter in my book about importance sampling which is way better. **Conditional MC**: very well explained in the paper Low discrepancy sequences. (Quasi-Monte Carlo techniques). It’s not explained well in the paper, but you can look it up and read better papers. Briefly: when you start generating one dimensional random variables, we’ll say for a uniform distribution, we can use testers to see if the numbers are actually uniform. Then you can generate pairs, two at a time (X\_1, X\_2) and plot them. You should definitely do this experiment with a random number generator. It would not look uniform at all. Human mind when you say uniform, thinks that it’s perfectly spread out. The point of the quasi generator is to spread out on purpose, to get a perfect uniform distribution. It’s okay as an exhaustive search method. But if you need more points, you have to quadruple them to have the same spread everywhere. The more detail you want, the finer quasi becomes, and it gets a lot slower. But there are circumstances in which this is useful, and people in engineering that don’t understand randomness like this thing. **Chapter 5** is about estimating American options using Monte Carlo simulations. ## subjective probability George Calhoun sent an article to me in Nature Trump’s election is a subjectiv eprobabiolity, it matters what people’s perceptions are. A stock, TSLA. It’s been going down, so you’r wondering if it should keep going down or should it stabilize? And we don’t know because the stock is not the value of the company, it reflects the perception of people about the company. It’s the same as poker. Game theory is so close to probability. When people lose concentration, they react poorly. ## Generating Correlated Processes If you have multiple stock processes that you want to generate that are correlated, how do you do that? Each asset looks like $$dX\_t^i \= f(X\_t^i) dt \+ g(X\_t^i) dW\_t^i$$ Each is a separate stochastic process, but they don’t move their own way at random. This Brownian motion is correlated. If we take the notation $$X\_t$$ the vector of components, and same for f(x). This is the simpler case. You can have a more complicated thing. Each function could depend on all the other ones, it doesn’t have to be driven by just one variable. g(x) is a matrix dxd size, where dW is dx1, the vector of Brownian motion. When you do the Monte Carlo simulation, the Brownian motion is what you’re interested in simulating. We will assume that these are correlated. With this notation, we have $$dX\_t \= f(X\_t) dt \+ g(X\_t) dW\_t$$ It’s more complicated because it’s a collection of multiple integrals on dW, but we will write like this for conciseness. We have the covariance matrix $$\\text{Cov}(\\Delta W\_t) \= \\Sigma \\Delta t$$ This is Brownian motion. So what exactly is the covariance matrix? The components are the covariances of each of the components. $$\\begin{bmatrix} 1 & \\rho\_{12} & \\rho\_{13} & \\rho\_{1d} \\\\ \\\\ \\\\ \\\\ \\end{bmatrix}$$ Symmetric and positive definite matrix\! Positive is vector times matrix times vector transpose, which is a 1x1 number. Positive definite means that for a vector u, this is always greater than 0\. This is very important because you can take a vector multiply with vector which will give us the variance, which will be always positive. For example, take Heston $$dS\_t \= rS\_t dt \+ \\sqrt{Y\_t} S\_t dW\_t^1$$ $$dY\_t \= \\alpha(\\bar{Y} \- Y\_t) dt \+ \\sigma \\sqrt{Y\_t} dW^2\_t$$ If I want to simulate this, which is part of homework, how would you do this? Generate two random numbers which are correlated with ρ. For two, it’s very simple. We need $$X\_1, X\_2 \\sim N(0, 1)$$, such that $$\\text{Corr}(X\_1, X\_2) \= \\rho$$ Then we can multiply by √Δt and everything will work. We will start with uncorrelated Z\_1, Z\_2 N(0, 1\) I’ll take X\_1 \= Z\_1. Then, I’ll take X\_2 \= ρZ\_1. We can do this because the variance of Z\_1 is 1, and covariance of Z\_1 and Z\_2 \= 0; The combination has to have variance of 1, and if we’re combining linear combinations, then it’s normal. So total value is $$X\_2 \= \\rho Z\_1 \+ \\sqrt{1-\\rho^2} Z\_2$$ If you understand the principle, what do I do when I want to do three? X\_1, X\_2, X\_3 Take the same idea $$X\_1 \= Z\_1$$ $$X\_2 \= \\rho\_{12} Z\_1 \+ \\sqrt{1-\\rho\_{12}^2} Z\_2????$$ $$X\_3 \= \\rho\_{13}Z\_1????$$ It becomes tricky. So is there a method to do this? There is\! You should know where this is coming from We have a vector X which is a lot of Xs. We are interested in covariance, which is distinct from correlation. Our covariance matrix will be one diagonal, where we multiply by √Δt. The values on the diagonal of the matrix are the variance, and the rest are covariance. So how do I generate a vector with this covariance structure? Here is the idea. If X is a random vector, (more details you could talk about in 540\) with mean μ, componentwise for each element in vector, then $$\\text{Cov}(X) \= \\mathbb{E}\[(X-\\mu) (X-\\mu)^T\]$$ Take Y \= AX. A is a matrix, but it can be ANY DIMENSION nxd. We can transform 4 components into 15 components with different linear combinations. The basis of the whole method is this: $$\\text{Cov}(Y) \= \\mathbb{E}\[(Y \- \\mu\_Y) Y \- \\mu\_y)^T\]$$ And we know that the expectation of Y from linear combination si $$\\mathbb{E}\[Y\] \= A \\mathbb{E}\[X\]$$ Because of this, we get $$= \\mathbb{E}\[(AX \- A\\mu\_X)(AX \- A\\mu\_X)^T\]$$ And you can factor this. Remember that the order is very important. $$= \\mathbb{E}\[A(X \- \\mu\_X)(AX \- A\\mu\_X)^T\]$$ And we can also do this in the transpose $$= \\mathbb{E}\[A(X \- \\mu\_X)(X \- \\mu\_X)^T A^T\]$$ And the inner product is the covariance matrix\! So it ends up being this $$= A \\Sigma A^T$$ **CHOLESKY DECOMPOSITION** Take the Z vector if iid normals, which are N(0, I\_d). On the diagonal, you have 1 correlation. Find a matrix A such that $$AA^T \= \\Sigma$$, the desired covariance. The reason why you do this is because you have A, and you multiply it, and you get the thing you need. Cholesky uses eigenvalues and eigenvectors, but it does exactly this. ## Covariance vs Correlation How do you get from the covariance matrix to correlation? $$\\rho\_{ij} \= \\frac{\\sigma\_{ij}}{\\sigma\_i \\sigma\_j}$$ How do you go from this method to other methods and vice versa? In terms of matrix operations, it’s not complicated. If D is the diagonal matrix of the variances (just the diagonal), then the correlation matrix is this. Inverse is regular inverse because it’s just a diagonal. $$= \\sqrt{D}^{-1} \\times \\Sigma (\\sqrt{D}^{-1})^T$$ And the transpose is actually the same thing. And you do non-inverse to get from correlation to covariance. Not commutative, so be careful. ## Markov Chain Monte Carlo Left unsaid ## Vietnam and the Tariff Shock - URL: https://sharifhsn.dev/blog/linkedin-2025-04-14-vietnam-and-the-tariff-shock/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-04-14-vietnam-and-the-tariff-shock.json - Description: The biggest loser of the tariffs, the tragedy of the "miracle on the Mekong": Vietnam 🇻🇳 - Date: 2025-04-14 - Exact published timestamp: 2025-04-14T15:40:44.045Z - Topics: Tariffs & Trade, Trade Policy - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7317572525679398915/ The biggest loser of the tariffs, the tragedy of the "miracle on the Mekong": Vietnam 🇻🇳 When President Trump revealed the list of tariffs on "Liberation Day", one country stood out: Vietnam. They were hit by 46% tariffs, the second highest out of all of them. But to really understand the scale of this mistake here, you have to walk back in Vietnamese history. Vietnam was destroyed after its war with the United States. It was run as a socialist nation for a decade, with strict agricultural collectivization and no free market trade. This led, predictably, to mass famine and slow growth. It seemed that there was no hope left for this country. The Soviet Union was their only lifeline, and the economic support from them dried up within only a few years. Then, in 1986, as the Soviet Union was on the verge of collapse, the Vietnamese government passed the Đổi Mới reforms. They made sweeping changes to the organization of the economy, allowing private enterprise. Vietnam opened itself up to the world, encouraging foreign direct investment and trade with the West. The success of these reforms led Vietnam to grow its GDP at a rate of 4-6% in the 1990s. Vietnam's socialist policies led them to redistribute these gains across the nation, leading to massive reductions in poverty. Today, Vietnam is akin to a middle-income nation like Mexico or Turkey, a far cry from its war-ravaged, famine-stricken past. Despite the trauma of war, Vietnam has significantly warmed its relationships with the U.S. in the past decades. Indeed, they are far friendlier to us than they are to China. The U.S. has taken advantage of this and shifted much of our manufacturing from China to Vietnam, reducing our dependence on China. This would seem like a win-win scenario. Not according to Trump. He saw our close trade relationship with Vietnam as a threat to American manufacturing, and slapped them with massive tariffs as a result. All of those years of careful diplomacy and relationship strengthening between the two countries are circling the drain. Although the tariffs have been paused, nobody is sure what will happen next. Vietnam is in talks with China and have already agreed to some cooperation agreements. [1] This can only be bad news for the U.S. I visited Vietnam last year. It's a beautiful nation, with delicious food, warm people, and a fascinating culture. On the flight there and leaving, they played a beautiful song called "Hello, Vietnam" on the speakers [2]. It's sung by a child of Vietnamese expats, who wishes to see the country where she originated. It's an anthem for Vietnamese immigrants, and paints a beautiful picture of Vietnam from a Western perspective. Unfortunately, it seems like the U.S. has decided to say "Goodbye, Vietnam". [1] [https://lnkd.in/eeX6mP2B](https://lnkd.in/eeX6mP2B) [2] [https://lnkd.in/e653TUpg](https://lnkd.in/e653TUpg) ## The Big Short and the Boring Cause of the 2008 Crisis - URL: https://sharifhsn.dev/blog/linkedin-2025-04-12-the-big-short-and-the-boring-cause-of-the-2008-crisis/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-04-12-the-big-short-and-the-boring-cause-of-the-2008-crisis.json - Description: I rewatched The Big Short last night, and realized that they sensationalized a pretty boring reason that the 2008 crisis happened. - Date: 2025-04-12 - Exact published timestamp: 2025-04-12T18:25:14.194Z - Topics: Financial Markets, Housing - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7316889148353630210/ I rewatched The Big Short last night, and realized that they sensationalized a pretty boring reason that the 2008 crisis happened. When the housing bubble burst and the Great Recession occurred in 2008, everyone wanted to find the individuals to blame. The easiest targets were the Wall Street bankers who made the bad bet that people wouldn't default on their mortgages, even as riskier and riskier subprime loans were issued. The Big Short is the most popular movie depicting these events. It focused on the maverick investors who realized that there was a bubble and bet against the housing markets. The villains of this movie are the bankers who created and sold the "CDO" or "collateralized debt obligation", which is what caused the housing crisis to become an economic disaster. The movie explains the CDO like this: imagine a chef buys some fresh fish which doesn't sell. He can throw the unsold fish into a stew, which becomes a whole new thing and is now palatable. In this case, the unsold fish are poorly rated BBB bonds, the stew is the CDO (a portfolio of bonds), and the new palatability is a AAA rating. But why would ratings agencies rate a bunch of terrible bonds packaged together as AAA? The movie frames this as a corrupt racket. The main characters ask their colleague at S&P why these securities are being rated AAA, and she admits that if they don't rate them AAA, then the banks will go down the block to Moody's and get a better rating there. They have essentially become a ratings shop. The reality of the situation was that the ratings agencies were using a flawed model, the Gaussian copula. In essence, the model assumed that bond default probabilities were uncorrelated. Any good portfolio manager knows that the key to risk management is diversification. By packaging together uncorrelated assets, you can reduce the risk of all of them going down at once. So in theory, if you have a lot of uncorrelated BBB bonds, the likelihood of many of them defaulting is fairly low. You can see that in this graph. The probability of portfolio loss exceeding 10% is vanishingly small for a portfolio of uncorrelated bonds. As correlation increases, such a risk becomes more and more likely. In this case, the likelihood of bond default wasn't uncorrelated. Unscrupulous mortgage brokers were issuing subprime loans en masse, making it likely that all of these people would default on their mortgages. In the wake of the crisis, on top of stricter banking regulations by Dodd-Frank, ratings agencies have adopted more complex models that incorporate bond default correlation, meaning that they're unlikely to make the same mistakes. The narrative of a ratings shop might play well to a lay audience, but it doesn't fit with reality. Did you know about this flaw, or did you assume that The Big Short had it right? Let me know in the comments. ## Romania's Export-Driven Debt Strategy as a Warning - URL: https://sharifhsn.dev/blog/linkedin-2025-04-11-romania-s-export-driven-debt-strategy-as-a-warning/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-04-11-romania-s-export-driven-debt-strategy-as-a-warning.json - Description: Trump's policy of reducing the national debt through an export-driven economy has parallels to 1980s Romania—and it doesn't end well. - Date: 2025-04-11 - Exact published timestamp: 2025-04-11T15:38:23.906Z - Topics: Sovereign Debt, Fiscal Policy - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7316484774297038848/ Trump's policy of reducing the national debt through an export-driven economy has parallels to 1980s Romania—and it doesn't end well. President Nicolae Ceaușescu was elected to lead the Socialist Republic of Romania in 1965, and became a totalitarian dictator in 1967. Among other seemingly anti-Soviet policies, his decision not to invade Czechoslovakia in 1968 made him popular in the West. Because of this, they generously lent money to Romania for the purpose of industrialization in the 1970s. However, due to internal corruption and the energy crisis of the late 1970s, Romania was unable to pay its debts. Like many other developing countries, it requested a line of credit from the IMF to pay off its debts. Unusually, however, Romania refused to negotiate with its creditors to restructure their debt and have them take a haircut, as those other developing countries did. They resolved to pay off their national debt on schedule and as soon as possible. For this purpose, Romania cut imports significantly and expanded exports. This austerity policy led to significant reductions in standards of living, with food and energy shortages throughout the 1980s. The export-driven economy further exacerbated this. Take the example of a shoe factory, since Romania is known for producing the best shoes in the world. All the most high quality leather would be reserved for shoes that were designated for exports. Shoes that were sold in Romania itself only used the leather that wasn't fit for export. The economic crisis led to further corruption, so people would steal the best domestic leather off the lines, leading to even worse quality for Romanian shoes on the market. In the end, Romania eliminated its national debt in 1989. But that same year, the political unrest caused by the impending fall of the Soviet Union led to revolution in Romania. The Romanian Revolution was the only one in the Eastern Bloc to turn violent, ending in the conviction and execution of Ceaușescu and his wife. The conclusion seems to be that the sharp drop in standard of living due to austerity was what led to the more severe consequences for the leader who put them into place. What do you think? Is Romania’s example instructive for the US? (Special thanks to Dr. Ionut Florescu for the detail on austerity in 1980s Romania) ## China's Tariff Strategy and the One-Punch Proverb - URL: https://sharifhsn.dev/blog/linkedin-2025-04-10-china-s-tariff-strategy-and-the-one-punch-proverb/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-04-10-china-s-tariff-strategy-and-the-one-punch-proverb.json - Description: An old Chinese proverb contains the key to understanding their tariff strategy: 打得一拳开 免得百拳来 - Date: 2025-04-10 - Exact published timestamp: 2025-04-10T13:30:12.308Z - Topics: Tariffs & Trade, Trade Policy - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7316090125531111425/ An old Chinese proverb contains the key to understanding their tariff strategy: 打得一拳开 免得百拳来 "throw out one punch now to avoid a hundred punches in the future" What does it mean? By striking hard all at once, you can avoid the pain and tedium of striking many times later. When President Trump placed a pause on the Liberation Day tariffs, the entire world breathed a sigh of relief. He promised that the tariffs would go into effect only on countries who had retaliated, which at this point was only China. China responded by escalating its tariffs to 84%. China is the only country refusing to back down from Trump's threats. This will devastate both economies. Last year, China imported $143.5 billion of goods per year from the U.S., and the U.S. imported $438.9 billion of goods from China. So why is China doing this? Trump's tariff policy has been, to put it mildly, volatile. Nobody is really sure what he will do next, and whether he will stick to his promises. By forcing the issue and escalating now, China is forcing Trump to deal with the economic reality of the tariffs. And unlike China, the U.S. is a democracy, where the leaders must listen to the citizens, who are loudly responding to the tariffs. Although the stock market has rallied in response to the 90-day pause, if the proposed 125% tariffs actually go into effect on China, we can expect to see further dips in the market as businesses are crushed by expenses. If China had backed down, we might have seen months and months of additional uncertainty around tariffs, with Trump pulling them in and out based on his whims. But with the pain that is sure to come now, Trump will have to face facts about the consequences of his policies, and hopefully reconsider his beliefs. What do you think? Is China's policy the right way to go? (Special thanks to Ke Ren for teaching me this proverb, which has gone viral on Chinese social media) ## Structured Credit and CDOs - URL: https://sharifhsn.dev/blog/advanced-derivatives-week-11/ - Structured data: https://sharifhsn.dev/api/posts/advanced-derivatives-week-11.json - Description: Last day of class is Thursday, May 8\. - Date: 2025-04-10 - Exact published timestamp: 2025-04-10 - Topics: Credit, CDOs, Structured Credit, Tranches, Synthetic CDOs - Categories: Credit - Source: Advanced Derivatives - Source URL: None ## Final Last day of class is Thursday, May 8\. The final exam is going to be posted then, take-home exam. It’ll be posted like an assignment. Will be posted in evening around 6:30, you’ll have Friday and Saturday night to do this. Questions will be theoretical. ## Multi-names Last lectures we discussed the mechanics of single-name derivatives like the credit default swaps, and the calculation of the implied hazard rates. This is what is calculated in the assignments, along with default/survival probabilities. Now we will discuss the multi-name credit derivative, mainly the **CDO: Collateralized Debt Obligation**. These are securities whose payments are linked to the incidence of default of an underlying portfolio of credit risky assets. It’s based on not just one, but a portfolio of risky assets. The most interesting component of this is modeling the dependencies in the portfolio, and then calculating the fractional loss in the portfolio, and essentially pricing. So you have a basket of such credit-linked assets, then the question is, if you want to sell a portion of this portfolio, in shares, what should be that price? Basically you need to calculate the cash flows, as well as expected loss. Same idea as fair price of two legs as in CDS. Some of these instruments are more complex, but in general, we can consider asset-backed securities. They are formed from loans, bonds, mortgages, and so forth. Usually, the income from the assets is tranched. A MBS has claims on money generated from these pools, cash flows that these mortgages generate from regular payments. If you aggregate these together, you get cash flows. These will be distributed to the owners of the pool. MBS that are created by adding many such mortgages and selling shares on the revolving pool. For example, such security is secured by the underlying mortgages. This kind of instrument is obviously more convenient to put in a portfolio than a single-name. If you want a single bond, you are dependent on the probability of default on the counterparty. If you have two bonds, then you have two probabilities of default. You might have probability p of default, and the probability of both defaulting is p1\*p2 (absent some dependencies between bond defaults). It’s an expectation that helps in constructing such pools. We can calculate under different assumptions. Let’s say we assume that you have no correlations for them, that would be a simple formula. You would want to get a distribution, the probability that you have 1, 2, n names default. If you have the probability of default, if you assume that’s a homogeneous portfolio where all bonds have same probability of default. Then you would enumerate the number of ways that you can have 3 bonds default in the portfolio, which forms the distribution. Some assumptions, it helps in some way to have diversification. But many times, these mortgage bonds wouldn’t sell because their credit rating was too bad, like DDD. You can have CMOs, which consist of multiple pools of securities, which are **tranches** (or slices). The idea is that even if it’s a homogeneous universe of very low-rated loans, like BBB, if you have the securitization of this and construct the CMO and CDO in such a way that they would have AAA rating. We will look at such a mechanism. ## Waterfall In the case of a CMO/CDO (depends what our underlying is), the structure of instrument is usually a “waterfall”. It defines how the income is received and how it’s distributed to different tranches. So, slide 5 with diagram gives us a simplification of asset backed securities. The underlying assets (bonds) are collateralized into a CDO. Let’s say we have such pool of assets. These assets do not necessarily have the same probability of default or principal. But if you aggregate these into a special purpose vehicle (**SPV**), the purpose is to generate tranches. These tranches, you have typically a **senior**, **mezzanine**, and **equity** tranche. Sometimes we will have more specifications. The principal from the original pool is split into these tranches. And you have a target return, a promised return under the assumptions that we have no default. Typically, the principal for the senior tranche is much larger than the mezzanine and the equity tranche. Basically, the waterfall, which is also called the structure subordination, this describes the scheduled coupon and principal payments from different securities. It also describes the losses. The arrangement of this pool will generate cash flows. This cash flow is used first to pay the senior tranche. Then after the senior tranche is completely paid, the cash flows go to mezzanine, then equity. So as long as you have no defaults, this system works. But when you have a loss, the loss is first taken by the equity tranche, then the mezzanine, and then the senior tranche. This structure of subordination or waterfall creates advantages and disadvantages for the equity tranche. This tranche is far riskier than the other tranches. The idea behind the whole mechanism is to have many market participants. The senior tranche has AAA rating. The rating agencies will rate these particular tranches. The magic of it was to take something rated BBB and turn it into something rated AAA. You could take the mezzanine tranche from multiple CDOs and get a new CDO from that. Things got kind of complicated. ## 2008 Along with the increase in prices of homes, the MBS market reached $4 trillion. Wall Street issued something like $700 billion in CDOs. This is a kind of interesting instrument. There were many problems. The mortgages given in the first place had various loose criteria. People would get houses without giving a down payment. The rating agencies also had a problem. The modeling of the dependencies of the names made the assumption that these were independent. It’s easy to calculate, but typically these names are not independent. You can argue that certain regions have high correlations, natural causes and such, so they are correlated more locally than all over the United States. During the crisis, these defaults were kind of correlated, so more defaults generated more defaults. They tended to use Gaussian models, which do not have tail dependence, which was the main issue. This kind of behavior wasn’t modeled, this possibility of loss. ## Synthetic CDOs A cash CDO is an asset backed security **ABS** where the underlying assets are debt obligations. The sense is that you own the underlying assets in the portfolio. The long position in the corporate bond is similar to a short position in the CDS. So it’s a similar risk. So another way to construct CDOs involves forming a similar structure. Instead of being long in the company bond, you can be short CDS. From a risk perspective it’s similar. The originator of a **synthetic CDO** can select a portfolio of companies that have a certain maturity structure, and look at the CDS for the companies. The notional is the total notional of the CDS contracts. Let’s say we have this example. We have an equity tranche which is responsible for losses in the underlying CDS until they reach 5% of the total notional principal. Let’s say this earns 1000 bp spread. The mezzanine tranche is responsible for losses between 5% and 20% (200 bp spread) and senior tranche is \>20% with 10bp spread. In case of default, when you have such a default, the idea is that the income is paid on the remaining tranche principal. If losses reach 8% of the notional principal, then the equity tranche is wiped out, and 3% is taken from the mezzanine tranche. So tranche 2 earns the promised 200 bps spread on 80% of its principal, because they have incurred losses. The promised return is applied to the surviving principal of that particular tranche. ## Single Tranche Trading You can trade tranches of portfolios of CDSs without actually forming the portfolio. Cash flows are calculated in the same way as if the portfolios were formed. We can discuss a little bit of the description of the waterfall for the single tranche CDO. In general we have the buyer of protection on the tranche, and the seller of the protection. The portfolio of short CDS positions is going to be used as a reference point which defines the cash flows between these two sides. This portfolio is not created just referenced. The buyer will pay the tranche spread to the seller, and the seller pays the amount that corresponds to the losses in the reference CDS. We can discuss the pricing model. PAGE 2 NOTES We have a simple payoff function, and this function is dependent on the cumulative default and percentage loss on the portfolio. We define some quantities which map the tranche to the reference portfolio. PAGE 3 NOTES These are two extreme points. And we can see how fractional loss becomes a linear function. PAGE 4 NOTES Then the function becomes very simple. PAGE 5 NOTES What is the mechanics of the payments? It depends on the evolution of this function. PAGE 5 NOTES CONTINUED PAGE 6 NOTES What is happening with the protection leg? This is how you protect yourself from losses on the CDS. PAGE 7 NOTES If there’s no change in the value, then there’s no payments. Typically we want to run Monte Carlo simulations. The problem is that you have one default, and then you get all these different cases. So you have to simulate very large numbers. We will look at these simulations in the second part. PAGE 8 NOTES The percentage loss is not just 1/Nc, because it may be reduced by recovery. PAGE 8 NOTES CONTINUED The loss is dependent on the detachment point. The cumulative losses in the reference portfolio are going to be greater than than in the tranche. Lecture\_8.pdf textbook image: This is an illustration of a possible realization. This one has some characteristics. In this case the reference portfolio consists of 100 credits with a FV exposure of $10M. That’s the reference portfolio. Then you can define a tranche, which has the FV of $30M, and the contractual spread is 250 bp. The attachment point is 3%, width is 4%, and a maturity of five years. Then you have a simulation for a possible example. Let’s look at some characteristics. You have a payment schedule which is quarterly. We can do some calculations. For example, we have one loss in Jun 2007 from the reference portfolio. So how does this translate? This is cumulative percentage loss in the reference portfolio. PAGE 9 NOTES Here we assume that it’s a homogeneous portfolio. This means we have 5 defaults in the portfolio before the tranche will begin to incur a loss. Let’s look at when we have 5 defaults in Jun 2009\. What is the tranche loss? This is going to be a linear function. This means if you get one such default, you get X% loss in the tranche. The tranche notional will reduce. There will also be a loss in payment by the protection seller. PAGE 10 NOTES You have two sides. The protection buyer will pay the tranche spread to the seller. They will get protection for a particular principal, which covers a particular cumulative default loss in that portfolio, a certain interval. In order to get protection for the losses. The losses are constructed out of the reference portfolio. So the payments are done for protection, which are the spread, which are specific to tranche. In the case of that protection. The net flow is from the point of the view of the seller. If you have positive cash flows, and defaults and then you have to make payments. ## Senior Tranche This works very well for the equity and mezzanine tranche. The senior tranche requires a small modification and some care. PAGE 10 NOTES CONTINUED Then you can calculate the tranche loss given a detachment point of 100% i.e. full loss. PAGE 11 NOTES This is the cumulative default loss in the portfolio. This doesn’t make sense because if the reference portfolio is wiped out, then the tranche should be wiped out. Then the senior tranche investors continue to receive payment in 44% of the tranche value. There is a solution to this: PAGE 11 NOTES CONTINUED In which case the fractional loss of such tranche is going to be map onto the maximum loss in the cumulative loss in the reference portfolio ## Correlation We are interested in the general portfolio distribution, and such portfolio loss distribution is going to indicate the probability of certain losses in the future. PAGE 12 NOTES It’s important to estimate this. We need some information about the portfolio loss distribution at different horizons, 1 default, 2 default, n defaults, 1 year, 5 years, n years etc. So this density is an important quantity. PAGE 12 NOTES CONTINUED Here we can assume that for different maturities, the recovery rates may be different. Nc is the number of names. This is the cumulative fractional loss in the portfolio. The important part is that we have a discrete number of names (credit derivatives) in the portfolio. Furthermore, such loss L(T) are weighted. We can do some simplifications, (in some cases you cannot do this and you have to do simulations), if we make some assumptions. PAGE 13 NOTES Recovery rate R\_I is a constant, but may not be known in advance. The expectation of the indicator function, that there is going to be a default, is the same as \[1-Q\_i(0, T)\]. The variance and the shape of distribution may be independent from the correlation, but the expectation is independent. We can write PAGE 13 NOTES CONTINUED In this figure Lecture\_8.pdf portfolio loss distribution figure. This figure shows the portfolio loss distribution, which is implied by three different values of correlation. This is generated by using the Gaussian copula model which we will discuss, essentially a multinormal distribution. Here we can look at loss distribution, other distribution of such loss function. We look at linked default between default correlation and single tranche CDO. As we saw, loss of portfolio is independent of default correlation. The expected value of the loss distribution is going to be the center at 5%, that makes sense. We illustrate when there is no zero correlation, and also when there is medium correlation and high correlation. What can we see? When the correlation is zero, the names do not default together. The portfolio loss is between 0% and 100%. If you have a senior tranche with an attachment point of 10%, you have a very low probability that such tranche is affected. In a high correlation environment, then the credits have a high probability of surviving default together. You have a high probability of losses exceeding 10%, so the senior tranche would incur loss. The conclusion of this, if you are a senior investor, and you are in an environment where you have very low correlation, then names should be as independent as possible. But typically, the institutions and financial firms that construct CDOs, they have an expected loss. If you hold such equity tranche, you might have a better chance of surviving this high correlation environment. ## Valuation of Tranches of Synthetic CDOs and Basket CDSs SLIDE 11 We will look at some copula functions that follow this. SLIDE 12 If we have a homogeneous portfolio, then… In this context Q is the cumulative default probability, not survival. We will discuss this first formula a little later when we discuss factor models. We assume we have some probability of default. But, this would be for the names that are not correlated. We can impose some dependency between the names in the portfolio. We will look at factor models, starting with single factor, then looking at multiple factors, and generalizing with copula models. Next week we will discuss this in more depth. ## Single Factor Models Lecture\_9.pdf We will generalize with a single Gaussian distribution, and then make more complicated. This model has flexibility when modeling such portfolio. PAGE 14 NOTES ## The Gaussian Latent Variable Model We will not go through the model, just introduce the quantities. This is a factor model. We have a random variable A\_i associated with credit i, each credit in the portfolio. And we assume that A\_i \~ N(0, 1). This particular model, we define that a default occurs when before time T if the value of A\_i is less than a time dependent threshold, which we will soon define. Each name in the portfolio is associated with a normal random variable, and we assume when we have a default. PAGE 15 NOTES The probability of default before T is the same as the probability that the random variable is smaller than or equal to this time dependent threshold C\_i(T). This A\_i is normally distributed. So this probability can be defined through the normal CDF. The probability of default can also be defined as 1 \- the survival probability (for credit i) up to T. If you have the CDS available on this instrument, that is another means of estimating the probability of default. This is a model that can be calibrated to the market. PAGE 16 NOTES Then we can get the time-dependent threshold. By having such means to calibrate, we can get the time-dependent threshold from the survival probability. SLIDE 3 In this particular table, what do we have? We have a relationship between the hazard rates and the survival probabilities, as previously discussed. PAGE 16 NOTES CONTINUED If you have such implied hazard rates, you can calculate the survival probability. This is very useful when you can use the CDS . For a small hazard rate, it means a higher survival probability, and vice versa. When survival probability is close to 1, 0.999, then the value of the threshold is the inverse cdf of 1 \- this value. This value is very small. So this is going to be on the left tail. The value in the threshold represents this. As you increase or decrease survival probability, this threshold changes. We have S\_i normally distributed, if it’s smaller than the time dependent threshold, it has a variable probability. And it’s based on the real market. We can take time present to the future, and it changes, because the survival probability does as well. If the survival probability is 0.5, then the threshold should be 0\. This is something that you can calibrate on the market. PAGE 16 NOTES CONTINUED REMARK Because A\_i has no dynamics and is not observable, it is latent, hence the name. SLIDE 4 We can see the relationship between A and the time-dependent threshold. It’s calibrated to the market, either through a single or multiple hazard rates. Either way you should be able to construct it. A\_i is fixed. At time tau\_1, A\_i hits the threshold and defaults. For each realization of S\_i, you get the time of default. Which means that you can get a distribution of times of default, expectation, and confidence. This is the relationship between S\_i and default. PAGE 17 NOTES For each value we simulate from A\_i, we get a time of default tau\_i. The algorithm is very simple, we calculate A\_i and calculate times of default, so we get the mean and distribution. This is a building block. Next week we will look at correlation. These names in the portfolio no longer need to have the same probability of default, each of them is dependent on its own hazard rate, and we can assume we have different nominal values, principal values, and we can also model initially a one-factor model, which has a single factor and also an idiosyncratic component, and then it’s about exposure to this factor. This offers a lot of flexibility in modeling. These models are about modeling the default time. PAGE 18 NOTES The idea is that you generate and calculate such default times. This will just give you the distribution of times of default. If the names are independent, you can generate all these times of default, and then you can get a distribution up to time T. This will have a vector. If you have 10k times, you can run this for all the credits. For time \= 1 year, what is the probability of having 1, 2, n defaults. Then you can calculate expected loss based on the principal. And this is calibrated on the market’s hazard rates. ## Next Week Will be remote, most likely. ## Why Crashing the Economy Won't Work - URL: https://sharifhsn.dev/blog/linkedin-2025-04-09-why-crashing-the-economy-won-t-work/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-04-09-why-crashing-the-economy-won-t-work.json - Description: The secret plan behind crashing the economy—and why it can’t work 😴 - Date: 2025-04-09 - Exact published timestamp: 2025-04-09T01:35:45.991Z - Topics: Macroeconomics, Monetary Policy - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7315547943300681728/ The secret plan behind crashing the economy—and why it can’t work 😴 The anticipation of and uncertainty around Trump’s tariffs have caused the stock market to plunge rapidly since Wednesday. Most people would view this as a bad thing. But it’s possible that this is the explicit goal of the Trump administration. The national debt is the source of much criticism from the administration. The purpose of DOGE is to cut the federal government of “waste, fraud, and abuse”, which combined with tariff revenue, is purported to reduce our national debt. So far, these claims are illusory. The national debt itself is not a big deal, but the interest payments are. They have ballooned in recent years, and much of the treasury debt is due to be rolled over at higher interest rates than when the bonds were initially issued. These interest rates are related to the yield of the treasury bonds. The yield is inversely related to the price, since the actual payoff of the bond is fixed. When bonds are in high demand, the price rises, and the yields fall. If the yields are low when the bonds are refinanced, we could see hundreds of billions in savings on interest payments. What might cause bonds to be in high demand? As a safe harbor, people want to buy them when they are reducing risk. When the stock market tanks, bonds should see higher demand as investors flee from equities to safer shores. So that’s the idea: crash the stock market, increase demand for bonds, reducing yields, then refinance our debt at the lower interest rates. There are some significant problems with this theory, however. For one, it’s not coming true. Although 10 year yields briefly dipped right after Liberation Day, they have risen back to previous levels. China also plays a strong role here. They hold a giant amount of treasury bonds. If they wanted to, they could sell those bonds en masse, massively increasing the supply, driving down the price and spiking yields. And with the way the US has been threatening them, it’s not an impossible scenario. There’s also the strong possibility that with political volatility in the US, bonds will no longer remain a safe investment. Equities and bonds in other regions like the EU are becoming more attractive, especially as countries like Germany begin to rev up their economies to stop relying on the US. Bottom line? Tariffs crashing the economy might be the real plan to reduce the national debt, but they won’t work compared to real political decisions like taxation and spending cuts. ## Monte Carlo and Variance Reduction - URL: https://sharifhsn.dev/blog/computational-methods-week-11/ - Structured data: https://sharifhsn.dev/api/posts/computational-methods-week-11.json - Description: We will cover this stuff in two different lectures. Today is Monte Carlo 1. - Date: 2025-04-08 - Exact published timestamp: 2025-04-08 - Topics: Computational Methods, Monte Carlo, Variance Reduction, Control Variates, Asian Options - Categories: Computational Methods - Source: Computational Methods in Quantitative Finance - Source URL: None ## Monte Carlo Methods We will cover this stuff in two different lectures. Today is Monte Carlo 1. The two theorems are the basis of the entire Monte Carlo simulation. **The Law of Large Numbers** **(LLN)** This says, very simply, given iid random variables X1, … Xn, mean μ, finite/bounded variance. If the variable has an infinite variance, it cannot converge because it goes all over the place. So that’s the one condition. Then There is zero probability I won’t converge. $$ \\bar{X} = \\frac{\\sum_{i=1}^n X_i}{n} $$ $$ \\mu = \\mathbb{E}[X] $$ How does it converge, though? It’s one thing to converge slowly, and another to be fast. This is governed by the **Central Limit Theorem** **(CLT)** that you also learn about in probability and statistics. This says that if you have the same situation X1…Xn iid, mean μ variance σ^2 \< ∞. Bounded means I have unknown variable Xi, variance keeps going up, σ \* n. Technically its’ finite, but it keeps going up. then \\(\bar{X} \approx N\left(\mu, \frac{\sigma^2}{n}\right)\\) This itself is kind of nonsensical, so we would say that the distribution of this transformed random variable approaches a normal distribution. $$ \\frac{\\bar{X} - \\mu}{\\sigma / \\sqrt{n}} \\sim N(0, 1) $$ But the approximation helps us understand it better: If you have fifty distributions, and you look at the distribution of those fifty, it will look close to some normal distribution. If you increase it even more, it will shrink the variance, and it will get closer to your object the estimate. And this gives you a numerical measure of how close you are, it allows you to express a *confidence interval*. At 95% C.I. \\(\\bar{X} \\pm 1.96 \\tfrac{\\sigma}{\\sqrt{n}}\\) Now technically we don’t know σ. So in practice we use the sample stdev S, which is $$ S = \sqrt{\frac{1}{n-1} \sum_{i=1}^{n} (X_i - \bar{X})^2} $$ Then if n is small, we use t\_n-1 instead of N(0, 1\) But we’re never doing less than 10k simulations so this is irrelevant to us. n degrees of freedom with greater than 100 n, it goes Normal. What does this have to do with anything? This is the key idea. We’re going to obtain somehow values for my stock in the future. Based on those values in the future, I’m going to estimate the value of my derivative. I estimate one path, and I pretend that’ smy path in the future. If I know that’s the path in the future, I can use the terminal value and current values of the path to calculate the derivative, if I know the values of the path. That’s my observation X1. Then I’ll do another path, which will be X2. And i’ll obtain all those rvs, and I can estimate μ, the expected value of the derivative. I’m using this formula to estimate an expectation. It works every time for this purpose because of the LLN. The question is, how do we calculate these things, with mean μ? The **real problem** needs you to forecast what X\_T is in the future. **Step 1**, most important step: Hypothesize a model of evolution from X\_0 to X\_T I know where I am now, I’m at X\_0. I want to know what are my assumptions about the world that will lead me from now until T. THis is more complicated, because you need to place all these assumptions. Let’s say I generate a stock price, which we’ve been doing, I want to know what the price of the fixed income of the treasury bond is, which is repaid a year from now. The Treasury is AAA, never default. You know what the price is. Can I calculate the rate of this instrument that is issued for this price right now, then gives me $1 payoff at maturity. But in the Trump era, but we don’t know if it’s safe anymore. So we need to figure out what the risks are. Then you need to play in things that aren’t part of your model? How do you implement this? You can do jumps. Between now and next year, Trump will do something to collapse the market, which will be permanent not temporary. Last week was temporary, nobody cares. But a permanent thing where the economy is severely damaged, that makes it so I might not get this money I’m getting. Then you need to introduce jumps. That requires you to do Monte Carlo with jumps, which is on the homework. But this Step 1 is the most important part. If you are financial engineer, you need to be able to understand randomness and how it relates to real life. You need to relate real life to this Monte Carlo simulation. The methodology by itself is very straightforward, but you have to understand how it relates to the real world. **Step 2** Once you have the model, you match the model to data. In the simplest example, I look at a model for my stock price, using GBM. GBM has two parameters, drift and variance. I look at historical data and estimate a long term drift and variance. I will use those two parameters to generate data for the future. Then you have to estimate your parameters. Then you also have to identify a way to generate randomness. This is a very vague statement, but what I mean is if I have a model that I hypothesize, the noise is gamma. I need to have a programming language that generates random variables as gamma random variables. First you need to justify why you need gamma and not normal, etc. Then you can generate randomness. We have this high frequency trading simulator stuff, and we use these zero intelligence agents to interact with real teams that are doing the trading. We need to worry about, how do we initialize these agents? How much money do we give them? We did some research that if you start with homogenous agents vs heterogenous agents, heterogeneous is much closer to reality. So how do you create this heterogenous? We use the Dirichlet distribution and sample from it to initialize the wealth of the initial agents. We identify the characteristics that the simulation needs to have, then identify distributions and rvs that fit that characteristic. Once you’re done with steps 1 and 2, **Step 3** is really simple. Create values for X\_T (iid), and that’s it. **Example with Call Option** Using X\_T1, you can calculate the value of C1 as $$ C^1 = (X_T^1 - K)_+ e^{-rT} $$ Then you repeat this for all your samples. Then we average from LLN: $$ \\bar{X_T} = \\frac{\\sum_{i=1}^n X_T^i}{n} \\rightarrow \\mathbb{E}[X_T] $$ These are all matching the expectation of the stock price: $$ \\bar{C} = \\frac{\\sum_{i=1}^n C^i}{n} \\rightarrow \\mathbb{E}[(X_T - K)_+ e^{-rT}] $$ **Furthermore**, the CLT tells us how close we are. I can construct the confidence interval Students get really confused about this. There are two ns. There’s the confidence interval for the mean, which gets closer to 0\. But it also appears in the second one, because that one is for the actual random variable. Not C bar, which is an average, but the actual C random variable. You have the random variable which is the value of the call. If you have one path, you get an estimate, one estimate, one call value. That estimate, that rv, is going to have a mean, that will be the true value of the call, and it will have a certain variance, which is probably huge. The way you estimate that, you take a sample, maybe 50, and then maybe estimate this quantity. It will get closer to the stdev. I have my variability, which could be $200. My original option, if I do the correct path, I’m within $200 of the real value. The more I have the n in Sdc, it will go to that 200 value, not 0\. This is plain vanilla Monte Carlo. What do we do with this? This works really well for European type options. But it’s tricky for American. There is a method called least squares Monte Carlo on pg 210-216. It does some kind of tree method, with the path, creates a distribution of the future based on where each path is. It’s not really Monte Carlo. It also works for Barrier Options. I mention this because we discussed them. The path needs to decide if it goes above or stays within a threshold. Asian options also work. The payoff is determined by taking an average over the lifetime of an option. If it expires in one month, and it’s calculated daily, you take daily stock values, and then the value of the stock is an average. Or you could have an average vs a terminal value, there’s different versions. But in general, one of the terms is the average value of the stock price. Monte Carlo can be used for ANYTHING. Nothing is financial inherently. If I want to examine deaths from a disesase because Trump took us out of WHO, you can make an estimate of the contagion. This X\_0 could be the number of infected people currently. Then you could have some kind of dynamic of how the disease spreads, and then X\_T is the number of people infected at time T. THen you can simulate multiple things, and then calculate all sorts of derivative prices, with some “payoff” of how much money I pay in hospitals, based on how many people die of disease. Monte Carlo is used everywhere, particularly in biology. ## Monte Carlo for SDEs Everything in our areas uses SDE. This is on MF pg 273\. I’ve been saying pages because it was written by two people independently. Monte carlo is written by Mariani, and Monte Carlo written by Florescu. This is a pretty general dynamic, you can multivariate, multidimensional. In our homework, we have to work out for Bonus how to apply this Monte Carlo to a multidimensional process (Heston). $$ dX_t = \\alpha(t, X_t) dt + \\beta (t, X_t) dW_t $$ This is odne in the book so you can look it up. How do you develop a Monte Carlo skill from SDE? You can assume that X\_0 is some fixed point. The first thing is to remember that this is the notation, the SDE doesn’t actually exist. The actual SDE is $$ X_t - X_0 = \\int_0^t \\alpha(s, X_s) ds + \\int_0^t \\beta(s, X_s) dW_s $$ ### Euler Discretization **Euler** is a very old mathematician. There was no SDE when he was living. His method worked to discretize regular integrals, so we are using the same exact method, which is Euler’s method. Euler discretization. We will take \[0, T\] divided into M. Typically you generate millions of paths. The number of intervals you create has nothing to do with number of paths. Good estimation requires number of paths. But for this thing it doesn’t matter. You use whatever frequency you want. Estimate a month call option. Now you have to decide, do I want to generate every day? The time interval will be 1/252 or 1/365, and generate 25 or 30 for a month, but it’s your choice, and it doesn’t make a big difference. The actual diffusion of your rv will be the same. It will matter if for example you estimate a barrier option. It’s about passing a threshold. If you have a lot of frequency, your rv might cross the barrier and you don’t know it. We have m intervals and m \+ 1 points, t\_0, t\_1, t\_m \= T. X\_0 will start at fixed x\_0 Then our process is: $$ X_i = X_{i-1} + \\alpha(t_{i-1}, X_{t_{i-1}}) \\Delta t_i + \\beta(t_{i-1}, X_{t_{i-1}}) \\Delta W_i $$ IT does not really need to be equally spaced, it works with everything. You can take day 1, 2, 14\. You just need to be careful about time between observations. The functions α and β is known. You choose the times. The only unknown noise is W. We know that increments are N(0, Δt), so you can calculate this as Z \* √Δt where Z is N(0, 1). 99% of the errors are coming from the student using the N(0,1) directly for W, and it’s too large. Always use the left hand point that you have. This is one path, you store the final value. But if you’re pricing the barrier option, you don’t just need to know this, you need to know path. The problem he gave us: ## Monte Carlo with Jumps Jumps allow us to introduce events that are unpredictable. You need to know they’re going to happen, but I don’t know when. I know that Trump will introduce tariffs again, because he’s a moron. We know that this is going to happen, but we don’t know why. We need this compound Poisson process. We discussed two ways. Exponential times we accumulate, or in our case (much better), since we have 0 to T and we know the time interval, we create the Poisson random variables, generated, with λ \* t as the expected number of events. Once you generate that, the Poisson tells you how many times our president will influence the financial market. Then you create these n uniform variables from 0 to T, and sort them, and those are the actual generated times when he says the stupid things. Then how are the stupid things going to influence the markets? Last week, he said tariffs on Wednesday, and markets went down. I am not watching the market, so tell me when to buy. NVDA is overpriced, TSLA will go down, I can’t make up my mind. This is being an investment banker or trader, not FE. We have to **Step 1** Identify our model. $$ dX_t = \\alpha(t, X_t) dt + \\beta(t, X_t) dW_t + dY_t $$ This can be ANYTHING, our choice, could be Black-Scholes, or something else. You just have to fit it to real data. Where Y\_t is a compound Poisson process. The sum of these jumps, where the time of these jumps happens according to an actual Poisson process. We know how to generate the α β part, same methodology. The only question is, what do I do with the jumps? I’m going to have to generate for each path, a set of jumps, and then add them to the price process. Generate Poisson(λT) \-\> k (every time a different k) Once I know k, I generate T\_1…T\_k times of the jumps for the Y\_t process. And these are in the interval \[0, T\]. If you forgot how to do this, you take each T\_i \~ Uniform(0, T) and sort them. What about jump amounts? Y\_1… Y\_k the size of the jumps It depends how you model your stochastic process. Let’s say I’m modeling Donald Trump. Maybe he says something about tariffs, he goes down. So there’s no point generating a rv that goes up, I know it goes down. So I will generate it on a distribution of negative numbers, uniform(-0.5, \-1). Let’s say the second jump will be positive, so do something like that. You need to adapt it to your case. But let’s just say for simplicity it’s iid. Homework says you should use the Normal distribution. Identify the intervals Δt which contain T\_1.. T\_k It’s possible that all the jumps happen really close to each other, in the same interval. I will just add all the jumps in the interval together. The jumps happen in the interval. Whenever you have the jump, you add the jump. If not, you don’t care. The path In the example in the homework, it uses returns, so it’s going to be logarithm S\_t plus other stuff, so be careful. Write down the math before you program anything. You might not be sure whether to multiply or to add. This is the basic idea behind Monte Carlo. Next week, he’s going to give us a better Monte Carlo method One of the other problems is a variance deduction technique. Euler Milstein is a better approximation than this one, but it only works for homogenous processes, which are not time-varying i.e. don’t contain x in α or β. ## Variance Reduction Imagine I’m pricing an option, and I look at the price I’m getting, and I get a 95% CI which is $10 wide, with n \= 100 paths. If you look at the difference, it’s \\(2 \* 1.96 (because it’s margin of error) \* \\frac{\\hat{\\sigma}}{100} \= 10\\) This is about $10. And then my boss tells me the price of the option is $3, to give him a better estimate. So how many paths do I need to do to get this to $3? Sigma hat is going to change, but that sigma hat converges to the true σ, so it’s supposed to be kind of close. If I divide by 10, I get $$2 \\cdot 1.96 \\frac{\\hat{\\sigma}}{\\sqrt{10000}} \= 1$$ I get $1 if I increase my sample size to 10000 paths. So then I go to my boss and he tells me the price is between $2 and $3. DId you forget that these options are sold in multiples of 100? It’s $200 to $300. This is way too wide a margin\!\!\! I want to get within 10 cents. So I have to multiply by 100, and now I get to 1 million. Now that’s a lot of paths\! And then if I want 0.01, which is within $1 of the actual contract, I need 100M paths. And it’s because of this square root. So what do we do? Variance Reduction techniques. Paul Glasserman is the expert at this. If you want to read more about this. He has this classic book Monte Carlo Methods in Financial Engineering (2003). Also a paper Monte Carlo methods for security pricing, 50 pages, but it does what the book does in a condensed way. This was written in 1997, so it’s old, but a lot of techniques still used today. Broadie sucks, by the way. This paper: [Paper](https://d1wqtxts1xzle7.cloudfront.net/97653387/BBG_jedc96-libre.pdf?1674420173=&response-content-disposition=inline%3B+filename%3DMonte_Carlo_methods_for_security_pricing.pdf&Expires=1744764322&Signature=I3cZHcahNRTM0-xMOo-cU96LFGfP51D-a2UMwsuXvCoBMTWMQw9Odh9DYY3sWeRIwJAHwmna-x3H3HiWMi~yebvujsqdyv2a2cWBZhJIT6gv06dt1iLOYfQmqfQI6y6PEUAtmFrZ657C493Y1Ac98Cl3lDtgpL-RnGNuXaVaGZ3ImTx8PiGr0uHM~mKZkL-zdyijEHrrks-9FYwi64-2GW1u7r3~J7NrkMH29zevcXT8pNy9yoqDvUntfJHuJ45YxzQY3Kq4u1D3sunt4SGVIrxqRNJ8WKvk8ZwiSe-uAO~B6317UFO-E2--5gWfee1fQ2glTOzNqH9Aigx343VK3A__&Key-Pair-Id=APKAJLOHF5GGSLRBV4ZA) Having more n is a better approximation, but it is not the solution, because it gets harder to reduce the confidence interval. So can we calculate σ hat smartly? The quantile doesn’t change (1.96), n doesn’t change, so how do I get a better σ hat. That’s wy these are considered Variance reduction techniques. ### Antithetic Variate Th doesn’t exist in Latin languages. English is confusing, sometimes you pronounce it differently. It’s the same writing but it’s different\! Always with the tongue in your mouth. Like the Sylvester from Looney Tunes, apotheosis of using the th sound. How does this work? This antithetic variate is based on a very simple idea. If Z \~ N(0, 1), then \-Z is also N(0, 1). That’s the entire idea. So how does this work? The problem with this path generation, what takes you time? The time generating the random variables. The idea is to generate the path with m increments. Z1, Z\_2, Z\_m, N(0,1) rvs. You can create a path using these rvs, and also \-Z\_1, \-Z\_2, \-Z\_m. For the same amount of numbers, you calculate two paths, where the second one is created for free. Unfortunately, it is not the same as having 2m independent paths. Then my stdev would decrease by √2, old interval, divided by 1.41. But these paths are correlated. So the estimate for σ hat is going to grow. This is a worse estimate. The net result is still better, but it’s not √2 better. $$C \= \\frac{\\sum\_{i=1}^m C\_i \+ \\sum\_{i=1}^m C\_i^a}{2m}$$ Really simple to implement, easy to do, but not that impactful. You do it because it’s free. Then we will talk about the versatile method, which is more complicated. ### Sidebar: ChatGPT Investing I don’t know how to invest. I asked ChatGPT. Tell me which stocks have the largest return drop since April 2 (Liberation Day). What I have noticed is ChatGPT is completely useless. I can’t even calculate. I said give me a table of returns showing largest drop. He found the three journal articles that talk about this, limit to what other people have published. Interesting: how do the results of ChatGPT correlate with the market. ## Control Variates The idea of this, this is more complicated. There are other simpler methods. Antithetic is the simplest. Control variate takes advantage of certain relations between derivatives. If you know certain things about derivatives, you can use that to your advantage. This first example is on page 207 of MF, well-explained. Related to **delta hedging**. How does this work? At t \= 0, we, the option seller, receive C\_0, the premium. This is the option price. The idea of delta hedging is that I pay the payoff C\_T to the option buyer. (Only works for European options). If we held Δ, the derivative of call price with respect to stock, units of stocks at time any time t. Then we can replicate the option payoff C\_T. If I hold that Δ units of stocks, at time T, I’m going to end up with exactly with the payoff. If the payoff is 0, I have 0 units of stocks, if I have one, I get one. The problem of course is that you can’t own something that moves all the time. How do I own a derivative? I own 2 shares, then it goes to 1.5 shares. Everyone hates this because you have to buy stock when it goes up, and sell when it goes down. The fact is that we don’t buy and hold. We have to adjust this delta hedge all the time. So how does this work in practice? Say we have time t\_0 \= 0, t\_i \= iΔt, Δt \= T/m, where m is number of intervals I’m going to take these equally spaced, option expires in 30 days, take every day. I find this easier to explain my way than the way other people do it. ⬇️check the notes for this table **Time | Receive | Pay to hold Δ units of S\_t** If I do this instantaneously, this is supposed to be 0\. The Δ is matched exactly. If you make m reasonably high, it should be close to 0\. But I can’t really add them together. Because these are cash flows at certain times, and you have to multiply them by the discount factor. You could either discount them all to present day, or compound them to time T. Multiply by e-rΔt, e-r2Δt, … or ermΔt, er(m-1)Δt, … If we are compounded, then we cancel out most of these, and then we get If the time increment is really tiny, I get an error. Then we will rewrite this whole expression. $$C\_0 \= C\_T e^{-rT} \- \\sum\_{i=1}^{m-1} \\frac{\\partial C\_i}{\\partial S} (S\_{(i+1)\\Delta t} e^{-r\\Delta t} \- S\_i) e^{-i\\Delta t} \+ \\eta e^{-rT}$$ And we will say η is approximately 0, noise that is discounted. This expression says that my value for my option, equals, using my paths, the intermediate steps. I know the final value, and I subtract it by all these steps. The only thing I don’t know is Δ, the derivative of the option. If I knew that, I would know how to generate my path. This control variate gives me a way better estimate. The only problem is I don’t know is Δ… except I do\! For Black-Scholes, Δ \= N(d\_1) (for a call). This d1 you can calculate with current value of stock price, and you know what it is based on your path. The idea behind this control variate method, which is conceptually so much more complicated, once you derive it is much easier. You generate paths using the Euler method, as explained before. The larger the model variability, the worse prediction. The standard error, the sigma over error, if you have no variance reduction, antithetic is some reduction. The control variate has a huge reduction in standard error. And this is typically the difference between financial engineer and computer scientist. You need to know what you’re doing. Monte Carlo is really bad, UNLESS you know what to do. And this is a lot of derivation. If you don’t have Black-Scholes, you need to find the Δ somehow, and that might take a long time. The variance might look good, but it will take longer. ## Asian Options This is the Asian option calculation, the example from the Monte Carlo paper. What is an Asian option? It’s an option whose payoff depends on an average price throughout the lifetime of the option. It’s usually easier to solv things in continuous time, and not discrete time. The most general type of Asian option: There are two types: ### Arithmetic The payoff looks like this $$\\frac{1}{T} \\int\_0^T S\_t dt \- S\_T \- K)\_+$$ This is the most general type you can imagine. It doesn’t exist in practice. This is the one mathematician slike to solve because it’s simpler. The integral takes the entire stok price, which you assume is continuous, calculate an average here. It’s basically the area. What you’re doing is, you’re taking the side, and dividing it by T. You’re actually finding out the height such that the area of the rectangle is equivalent to the area under the curve (integral). That’s the arithmetic average. And, it is including S\_T and K. Typically it is only one of them. I wrote it in the most possible general way. You can get even more general. ### Geometric The payoff here is written like $$(e^{\\tfrac{1}{T}\\int\_0^T \\ln S\_t dt} \- S\_T \- K)\_+$$ This is called geometric, I’ll explain why. It’ll reduce to a geometic average on discrete formula. But this looks so ugly\! For this, we have an analytical formula. But for arithmetic, we don’t have a formula. The reason we have a formula is because ln S\_T is normal, which means that this is an mgf, and we have a formula that we can calculate with. In practice, none of this exists, it’s theoretical baloney. But in practice, everything that is solved is arithmetic. Nobody understands geometric average. What is that? ### In RL And how does it look like in practice? Monitor the stock price at certain times, from S\_t\_1… S\_t\_m. You’re going to observe the price every end of the day. This is defined int eh contract. The contract says, look at the average, it’s calculated using beginning of the day prices, the high during the day, low, average, etc. That’s a specification on the S\_t\_i. The arithmetic Asian: $$(\\frac{\\sum\_{i=1}^m S\_{t\_i}}{m} \- K)\_+$$ This is a call This is the same as the integral. You can take T/m as the integrand, you approximate LHS, and that’s what you get. T cancels, increment is T/m. For geometric, If you use the previous formulation, it is $$e^{\\tfrac{1}{m}\\sum\_{i=1}^m \\ln S\_{t\_i}} \- K)\_+$$ This is replacing the integral with an approximation, t cancels. THis is equal to, if you cancel out all of these sums of logs, you get $$= (\\sqrt{S\_{t\_i} \\ldots S\_{t\_m}} \- K)\_+$$ Except not square root, it’s mth root. This is a lot harder to price than using the initial formula. And geometric has a formula\! If you assume that the underlying is GBM, same as Black-Scholes. Arithmetic does not have ea formula. But in the market, there are zero contracts that are sold using the geometric average, because no one understands it. But everyone trades using the arithmetic Asian option. So how do you price the arithmetic, using the geometric? You use the control variate thing. First, we generate a path i, which is S\_t\_0, S\_t\_i, S\_t\_m you calculate the arithmetic premium A\_i from the same formula as before: $$(\\frac{\\sum\_{i=1}^m S\_{t\_i}}{m} \- K)\_+ e^{-rT}$$ You can also calculate G\_i the geometric premium from before: $$= (\\sqrt{S\_{t\_i} \\ldots S\_{t\_m}} \- K)\_+$$ I have three different ways to do this. One is one that I came up with, which you can’t find anywhere, and I don’t know if it works. Two is given by controlv ariates. You have all these A\_i and G\_i. It would make sense to do some kind of regression. I assume that there is a certain distance between A\_i and G\_i. A which is the true price, follows $$A \= \\alpha \+ \\beta G$$ This is a theory, but it must be true because these are two numbers, in general. What you do is a regression between A\_i and G\_i. $$A\_i \= \\alpha \+ \\beta G\_i \+ \\epsilon\_i$$ The idea is to minimize the noise, you’re fitting a regression so there is some noise. We can come up with three different estimates. Estimate 1: Don’t use any control anything. Just use A bar, just the average, regular Monte Carlo. Estimate 2: This is what control variate is supposed to do. I have a formula G. I don’t need any estimation for that. In that formula, which is a function f(S\_0, K, T, r, σ). I don’t need any path, so I know what that number is, without doing any approximation. I basically take A bar, and adjust it by the following adjustment. $$\\bar{A} \- \\hat{\\beta}(\\bar{G} \- G\_{\\text{formula}})$$ You can derive this, it’s not very complicated. That’s the other estimate. EXCEPT THAT, I don’t understand what’s wrong with #### Estimate 3 It would make sense that the relationship between A and G should be the same as the theoretical relationship between real A and real G. So why don’t we just take $$\\hat{\\alpha} \+ \\hat{\\beta} \* G\_{\\text{theory}}$$ These intrinsically depend on all the paths, so they contain information about the paths. Pretty sure I’m going to give you a problem for this, cook it up on the final. ## Misinformation in Volatile Markets - URL: https://sharifhsn.dev/blog/linkedin-2025-04-07-misinformation-in-volatile-markets/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-04-07-misinformation-in-volatile-markets.json - Description: Misinformation becomes especially dangerous in volatile markets... 😱 - Date: 2025-04-07 - Exact published timestamp: 2025-04-07T19:31:32.998Z - Topics: Financial Markets, Misinformation - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7315093897339101184/ Misinformation becomes especially dangerous in volatile markets... 😱 After President Trump's tariff announcement on "Liberation Day" last Wednesday, markets have roiled. The easiest way to see this is measuring VIX. This index averages the implied volatility of SPX options expiring in a month. It is known as the "fear gauge" because it goes up when markets are uncertain, and therefore fearful. VIX typically hovers around 20-30, but after Liberation Day, it has spiked to 40-50. Market participants are waiting with bated breath for any news on the tariffs, whether they will be permanent or temporary, or be exempted from certain markets. So when a tweet went viral from someone named "Walter Bloomberg" declaring that Trump was considering a 90-day pause on tariffs, the markets reacted immediately. The S&P 500 jumped up 300 points in just *ten minutes*, representing over $2.5 trillion in value. But this optimism wouldn't last. When the market realized that the tweet had no factual basis (it was a wild misinterpretation of Hassett's remarks on Fox News days earlier), the market tanked back to its earlier levels. And after the White House confirmed that the tweet was fake news, the S&P continued as it had before, albeit with more volatility than normal. It's a sharp reminder of how even the smartest traders on Wall Street can get taken in by seemingly-plausible BS on social media. If they can be fooled, so can you. Read more at Bloomberg: [https://lnkd.in/eBRP9jZg](https://lnkd.in/eBRP9jZg) ## Documentation Is the Key to Vibe Coding - URL: https://sharifhsn.dev/blog/linkedin-2025-04-05-documentation-is-the-key-to-vibe-coding/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-04-05-documentation-is-the-key-to-vibe-coding.json - Description: Documentation is the key to vibe coding success! - Date: 2025-04-05 - Exact published timestamp: 2025-04-05T19:46:48.390Z - Topics: AI, Software Engineering - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7314372961040289792/ Documentation is the key to vibe coding success! The use of LLM AI code assistants (or "copilots") like Github Copilot and Claude has exploded recently. Many non-programmers have taken advantage of this to begin coding their own projects without any past software engineering experience, just "coding on vibes". However, this is not as easy as it might sound. There's a lot of basic principles and best practices to software engineering that copilots don't steer you away from. If you make a design mistake, your copilot might double down on it and spit out a huge amount of code to solve a simple problem in a wrong, complex way. Ultimately, the only solution is to actually know how to code, how to use your language, and the library you're using. But there are some things you can do in the meantime to improve your output. The key to understanding here is that LLMs are input/output machines. The output is shaped heavily by the information and context you give it. If you just give the prompt "Make a full stack calculator app in React", it's going to have to make a lot of assumptions. Some, even most of those assumptions will be fine. But ultimately you won't know the difference, which will lead to pain later on when some of them aren't. Therefore, for any vibe coding project you're embarking, detail extensively in a document all of the project requirements. Any bit of information you have in your head related to the project, write down. If you feed that into your copilot, it will be able to reference that information when it's relevant and tailor its output to what you give it. This might seem daunting, especially if you're vibe coding something you don't fully understand yourself. But LLMs can help with that, too! Before you start actually coding, ask the LLM what information it thinks you can provide to help with its output. With just fifteen minutes of preparation and research, you can save yourself hours of debugging later. If you have any other vibe coding suggestions, please share in the comments! ## Rust 1.86: `get_disjoint_mut` - URL: https://sharifhsn.dev/blog/linkedin-2025-04-03-rust-1-86-get-disjoint-mut/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-04-03-rust-1-86-get-disjoint-mut.json - Description: Rust 1.86.0 has been released, with a very nice ergonomic feature to get around the borrow checker: get_disjoint_mut - Date: 2025-04-03 - Exact published timestamp: 2025-04-03T18:41:40.946Z - Topics: Rust, Programming Languages - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7313631796301099010/ Rust 1.86.0 has been released, with a very nice ergonomic feature to get around the borrow checker: `get_disjoint_mut` Rust's borrow checker is the heart of its promise of memory safety. On a basic level, it ensures the following - when you obtain an immutable reference to an object, no mutable reference can be made to that object while that reference exists - when you obtain a mutable reference to an object, that reference is the only reference that exists This creates a problem with slices, however. Imagine that you have a buffer, and you want to mutate only the last element, and pass around the rest of the buffer through immutable reference. You can take a mutable reference to the buffer to mutate the last element through `get_mut`, but you can't take an immutable reference to the rest of the buffer through `get`. Even though you know you're only mutating the last element, the borrow checker doesn't know that. This can require unergonomic reshuffling of memory access, or in the worst case, unnecessary cloning. The new `get_disjoint_mut` solves this problem. By returning mutable references to only the slices that you intend to mutate, you can have finer control over mutable memory access. Read more details in the release notes. [https://lnkd.in/e5Je6Mb3](https://lnkd.in/e5Je6Mb3) ## Credit Default Swap Valuation - URL: https://sharifhsn.dev/blog/advanced-derivatives-week-10/ - Structured data: https://sharifhsn.dev/api/posts/advanced-derivatives-week-10.json - Description: We discussed the mechanics of CDS a bit, and then started on the valuation. So let’s discuss the valuation in more details. - Date: 2025-04-03 - Exact published timestamp: 2025-04-03 - Topics: Credit, Credit Default Swaps, CDS Valuation, Premium Leg, Protection Leg - Categories: Credit - Source: Advanced Derivatives - Source URL: None ## Credit Default Swaps We discussed the mechanics of CDS a bit, and then started on the valuation. So let’s discuss the valuation in more details. We have two legs. The premium leg is the buyer, who buys protection against default for the underlying company. They have to make scheduled payments throughout the life of the CDS. So that if there is a credit event, there is a payment on the premium which has accrued. If there is a default in this period, you still have to pay the accrual. In case of default, the protection leg will purchase the underlying bond at face value. That’s the other leg. ## Valuation of the Premium Leg (O’Kane 6.5) So we want to determine the spread of the CDS. That represents the percentage of the principal that we have to pay to purchase the CDS. Let’s review a little bit what we did last time. We defined the present value of $1 to be paid at t\_m which cancels with zero recovery on default before t\_m This is a *risky* ZCB as opposed to riskless. $$ \\tilde{P}(t, t_n) = \\mathbb{E}[e^{-\\int_t^{t_n} r(s) ds} \\cdot \\mathbb{I}_{\\tau > t_n}] $$ Furthermore, if you assume independence between short rate process and the time of default τ, which is a reasonable assumption, then the expectation can be separated as a product because the joint expectation can be written in such a way. We can take the survival probability: $$ \\tilde{P}(t, t_n) = P(t, t_n) Q(t, t_n) $$ Then you can write the present value of premium leg as $$ S_0 \\sum_{n=1}^N \\Delta (t_{n-1}, t_n) Q(t, t_n) P(t, t_n) $$ where Q(t, T) is the survival probability at time t of the reference entity to T. The Δ(t, T) is the day count fraction between T and t. This is applied to some principal. This is the expected payment. ## Premium Accrual Another component of the premium leg is the premium accrued. The amount of premium accrued at default is *contingent*. The price today of $1 paid at default which occurs at \[s, s+ds\] is given by $$ P(t, s)[-dQ(t, s)] $$ I want to calculate the accrued premium for an infinitesimal amount, then integrate, because we don’t actually know when the default is going to happen. So what is such premium accrued? This is the expected such contingent payment, so it is equivalent to the expected default in that particular interval. So we’re going to get an amount with some probability of default, then we multiply the probability of default in that small interval. So that survival probability is decreasing, an infinitesimal change in Q will correspond to the probability of default between t and s. Then we discount such contingent payment back to time t. So then you start with CDS Spread S\_0, from the time of previous payment to time s (default). That is the amount accrued on expectation. $$ \\text{Accrual} = S_0 \\Delta (t_{n-1}, s) $$ Then the expected present value of premium accrued due to a default in \[s, s \+ ds\] in the nth premium period is… The amount is $$ \\text{Accrual} P(t, s) [-dQ(t, s)] $$ Note that default can happen any time during the coupon period. So the value of accrual is integrated. $$ S_0 \\int_{t_{n-1}}^{t_n} \\Delta(t_{n-1}, s) P(t, s)[-dQ(t, s)] $$ This is a quantity such that expected accrual in case of default between t\_n-1 and s. You can sum this over the life of the CDS. Then you will have a sum of such integrals, over each payment period. **This is the exact formula:** Sum over all the premium payment periods to calculate **the expected present value of the premium accrued** $$ PA = S_0 \\sum_{n=1}^N \\int_{t_{n-1}}^{t_n} \\Delta(t_{n-1}, s) P(t, s) [-dQ(t, s)] $$ **The present value of the premium leg** therefore becomes $$ PV_{\\text{premium}} = S_0 \\left[\\sum_{n=1}^N \\Delta(t_{n-1}, t_n) P(t, t_n) Q(t, t_n) + PA / S_0\\right] $$ ## Approximation of Protection Leg This is an exact expression. But there’s a problem with integration. You need the survival probability at each point. You can do an approximation, by valuing the function at the end of the interval and taking an average. Important assumption: the payment happens at the end, when default occurs. $$ PA/S_0 \\approx \\frac{1}{2} \\Delta(t_{n-1}, t_n) P(t, t_n) [Q(t, t_{n-1}) - Q(t, t_n)] $$ Basically, the change in survival probability represents the probability of default in that interval. Sometimes, it’s difficult to have information about the survival probability. If we look at bonds, we make the assumption that every year has the same probability of default. In terms of the formula, it’s exact, but the approximation is more useful. Here we assume that on average, default will happen halfway. ## Valuation of Protection Leg (O’Kane 6.6) The protection leg is a contingent payment of par minus recovery on the face value of credit following a credit event. We have an uncertain quantity which is paid at default. We have an assumption that we have a certain recovery rate, but we don’t know exactly how much can be recovered (depends on the restructuring process). The price of a security which pays an uncertain amount (φ(τ)) at time of default: $$ \\tilde{D}(t, T) = \\mathbb{E}[e^{-\\int_t^{\\tau} r(s) ds} \\cdot \\phi(\\tau) \\cdot \\mathbb{I}_{\\tau \\leq T} ] $$ This is a more general expression, but in general in the problems, we assume $$ \\mathbb{E}[\\phi(\\tau)] = 1 - R $$ If you recover nothing, the payment is the full principal. So we assume this is a fixed number, depending on the type of loan. This is independent of interest rates and default time. Then the present value of the protection leg is.. $$ PV_{\\text{protection}} = (1 - R) \\int_t^T P(t, s) [-dQ(t, s)] $$ You have one payment but you don’t know when you’ll get it. So you take the expectation. The infinitesimal here is the probability of default for each infinitesimal, then discount it. This is the general formula, which is complex to use. ## Approximation of Protection Leg One simpler approach is to discretize the time between t and T into k intervals: $$ \\epsilon = \\frac{T - t}{k} $$ Then we can turn this integral into a sum over discrete intervals. $$ PV_{\\text{protection}} = (1 - R) \\sum_{k=1}^K P(t, k\\epsilon) [Q(t, (k-1)\\epsilon) - Q(t, k\\epsilon)] $$ Another approximation: We know that P(t, T) is a monotonically decreasing function of T, so we can define bounds lower: $$ L = (1 - R) \\sum_{k=1}^K P(t, t_k) [Q(t, t_{k-1}) - Q(t, t_k)] $$ Basically, here we have a difference in survival probabilities, where one is discounted. The problem is that the cash flows have to be discounted, but right now it’s discounted at the end of the interval of kε. The discount depends on when the default happens, which can be at any time during the interval. $$ U = (1 - R) \\sum_{k=1}^K P(t, t_k) [Q(t, t_{k-1}) - Q(t, t_k)] $$ The lower bound is for large value of T, upper bound is for smaller value Then we take an average to get: $$ PV_{\\text{protection}} \\approx \\frac{1}{2} (L + U) $$ And then for a fair value of a CDS, based on no-arbitrage argument $$ PV_{\\text{premium}} = PV_{\\text{protection}} $$ **SPREAD FORMULA:** Then we can finally solve for the spread \\(S_0\\) $$ S_0 = \\frac{(1-R) \\sum_{k=1}^K [P(t, t_k) + P(t, t_{k-1})] [Q(t, t_{k-1}) - Q(t, t_k)]}{\\sum_{n=1}^N \\Delta(t_{n-1}, t_n) P(t, t_n)[Q(t, t_{n-1}) + Q(t, t_n)]} $$ Therefore, we need the full discount curve, the survival curve, this information. ## Example S \= spread We assume that payment occurs at the end of the year. We have the survival probability that the name survives, then we have an expected payment which is discounted to the present. Let’s do one calculation: If we have a static probability of default, probability of default in second period is Q\*(1-Q), probability of survival in second period is Q^2. Expected payment in 3rd year is Q^3 \* S. PV of expected payment is Q^3 \* S \* e^{-r \* T} The PV of accrual payment in event of default: The more general formula is on the average of default happens halfway, so our time of default is 1.5, 2.5, etc. If during third year, the probability of default (from table) is 0.0192. Then the coupon is halved, and the accrual payment is 0.5S. If it’s every month, then you divide that by 12 as well. Expected accrual payment at t \= 2.5 iii With the CDS rate on the market, you can make an inference about the implied probability of default. If mid market spread for a 5 year CDS is 100bps per year, then the conditional default probability is 1.61%. ## Implied Hazard Rates from CDS Spreads For the problem of estimation of default probabilities and corresponding hazard rates, we will used the so-called **JPMorgan model**. This is the credit curve, the “term structure” of probabilities of default. ## Review of Hazard Rates τ \= time to default CDF of probability of default is F(t) \= P(τ ≤ t) For the purpose of this, we will take survival probabilities, which are 1 \- F(t) The hazard rate is either h or λ by convention: The survival probability is equal to the $$ S(t) = e^{-\\int_0^t h(u) du} $$ If the hazard rate is constant, you have e-ht function. But this doesn’t have to be constant. Typically, we will assume it’s piecewise constant, depends on how we estimate it. If S(t) is differentiable, then $$ h(t) = -\\frac{d}{dt} \\ln S(t) = \\frac{F'(t)}{1-F(t)} $$ This is the relationship between the hazard rate, survival, and default probabilities. If you have a risky ZCB with zero recovery that pays $1 at T: $$ \\tilde{P}(0, T) = \\mathbb{E}[P(0, T) * \\mathbb{I}_{\\tau > T}] $$ where $$ P(0, T) = e^{-\\int_0^T r(u) du} $$ Then, survival probabilities are $$ S(t) = \\frac{\\tilde{P}(0, T)}{P(0, T)} $$ We can also note that the risky ZCB value must be less than the riskless ZCB.. average of hazard rates will be time 0 Sometimes it’s easier to calculate an average hazard rate for some maturity, then afterwards compute the piecewise constant hazard rate. We can determine h\_i from successive survival probabilities. We have the survival probability, but to express in terms of hazard rates. If you assume hazard rates are known ## Calibration and SDE Parameter Estimation - URL: https://sharifhsn.dev/blog/computational-methods-week-10/ - Structured data: https://sharifhsn.dev/api/posts/computational-methods-week-10.json - Description: Parameter estimation for stochastic differential equations. - Date: 2025-04-01 - Exact published timestamp: 2025-04-01 - Topics: Computational Methods, Calibration, Maximum Likelihood, Stochastic Differential Equations, Parameter Estimation - Categories: Computational Methods - Source: Computational Methods in Quantitative Finance - Source URL: None ## Plan for Today Parameter estimation for stochastic differential equations. ## Romania investments by the West after dictator Ceaușescu failed to invade Czech republic he stopped imports and increased exports romania makes the best shoes (dress shoes) they exported with the best leather, shit leather was for romanian shoes black market, people stole best shoes off line they had tons of money, but nothing to buy we had alcohol, that’s it 85 is some stupid stuff (1984) ## Calibration You observe some data. Most of the time, you observe the data in time. That is traditional for our domain, and many others. We say okay, I need to understand the dynamic of this data. In order to do this, we introduce randomness, because the randomness allows you to create data which is different. There is no way in hell you can predict the next data. I can try to predict the trend, and predict the variability around that trend. In our area of finance, it turns out that the stuff is so random, we have to go to these SDEs to understand anything. In that stochastic equation we have two terms, drift and diffusion. Drift is trend, diffusion is variability. Once I hypothesize a model, the next step is to find the parameters. For BS, estimate sigma and mu. You can calibrate, or estimate from path of process (more complicated). Calibration, you look at a derivative. A function of an underlying. We need financial derivatives. We have seen derivatives like calls and puts. We hypothesize a model, where you come up with a formula for the derivative prices. For example: $$ C(S, K, T, r, \\text{parameters:}, \\theta_1, \\theta_2, \\theta_3) $$ Let’s take the Ornsteinn-Uhlbeck model: This is a mean reverting process. $$ S_t = \\theta_1(\\theta_2 - S_t)dt + \\theta_3 dW_t $$ This is called the mean reverting Ornstein-Uhlenbeck model. How do you estimate this? Let’s hypothesize an option on this, with all the known parameters known. Given this, I’m going to solve a differential equation and get some sort of formula for the solution. Once you have that, you’re done. Then you’re basically done, because you have now a massive complex optimization problem. Let’s say you observe \\(C_1, C_2, \\ldots C_k\\) market option prices. Then you calibrate the model to these option prices. So you take $$ \\min_{\\theta_1, \\theta_2, \\theta_3} \\sum_{i=1}^k (C(S, K_i, T_i, r, \\theta_1, \\theta_2, \\theta_3) - C_i)^2 $$ (keeping in mind which are subscript i, different for each call option) You can change this. Let’s say you have twenty options, but you’re more interested in fitting out of the money options, you can weight each of them. Let’s say you want the four options like this, then you weight one fifth of each of those options, and one fifth of the rest. ‘ Now in practice, for instance, we talked in the beginning of class about local vol. That’s a calibration method, we calibrate models to the observed values. In practice, if you’re doing real quantitative analysis, you do have these complicate model which have a combination of fitting/calibration, and another method (estimation). ## Estimation This is a very classical statistical problem. We have the SDE: $$ dX_t = f(X_t, \\theta) dt + g(X_t, \\theta) dW_t $$ Then the O-U model looks like $$ f(x, \\theta) = \\theta_1 (\\theta_2 - x) $$ $$ g(x, \\theta) = \\theta_3 $$ Then for CIR, $$ f(\\theta) = \\theta_1(\\theta_2 - x) $$ $$ g(x) = \\theta_3 \\sqrt{x} $$ I want to estimate parameters based on data. We assume that we have observations $$ x_1, x_2, \\ldots x_n $$ at times $$ t_1, \\ldots, t_n $$ They could be same distance but it could be any times you want to observe this process. So how do you estimate the vector of parameters. In theory, you do this by **Maximum Likelihood Estimation (MLE)**. There’s also method of moments, and Bayesian methods Bayesian requires specific distributions, kind of complicated. Method of moments is kind of like calibration to observed moments. MLE is theoretically the most powerful one. How does it work? It’s really simple: It works with the density of these observations. If you’re observing the values of the process at \\(t_1, \\ldots t_n\\) You can form the joint density: \\(f(x_1, x_2 \\ldots x_n | \\theta)\\) This result depends on the vector of parameters, where θ is given. The method of maximum likelihood says that I’m actually observing x\_1, x\_n. The probability of observing those things should be the highest in those numbers. So let’s form a function $$ L(\\theta) = f(x_1, \\ldots x_n |\\theta) $$ In this expression, I’m going to plug in al lof the observations I know, and make this function of a function of the parameters, which I don’t know. Typically you take the log which is called the score function. $$ \\log L(\\theta) = l(\\theta) $$ Computer scientists came up with it, they sell better than mathematicians. If we know the joint density, then we can solve this problem. The problem is that we don’t know the joint density. It’s a really complicated expression. How do we do this? **The first trick**: we need $$ f(x_1, \\ldots x_n |\\theta) = f(x_n | x_1 \\ldots x_{n-1}, \\theta)f(x_1, \\ldots x_{n-1}|\\theta) $$ Joint divided by marginal. You can continue this up until the very first one. $$ = f(x_n|x_{n-1}, x_{n-2} \\ldots x_1, \\theta) f(x_{n-1}|x_{n-2} x_{n-1} \\ldots x_1 \\theta) \\ldots f(x_2|x_1, \\theta) f(x_1 | \\theta) $$ Remember there should be theta all over the place. These are all functions on one variable. These are all solutions to my stochastic process. Since x solves an SDE, then x is Markov. Which means that it’s the same as $$ x_n | x_{n-1}, \\theta $$ All I need is the previous value. So this is the generalized function. $$ x_{n-1} | x_{n-2}, \\theta $$ etc. So then the likelihood function is the product $$ L(\\theta) = \\prod_{i=2}^n f(x_i | x_{i-1}, \\theta) f(x_1 | \\theta) $$ This is actually not the same function. There’s another thing to be aware of. This is a **homogenous process**. Distribution does not depend on the time you are collecting it, only depends on the difference of the times. If you’re looking at a non-homogenous process, the distribution of today, tomorrow, and the next day are not the same as a year from today, next day from that, next day from that. Stationary is always the same, homogenous A non homogenous process would have a distribution of \\(X_1, X_2, X_3\\) vs \\(X_{100}, X_{200}, X_{300}\\), this is always going to be very different. \\(X_{101}, X_{201}, X_{301}\\). For a non homogenous process, these are also different. But for a homogenous process, this will be the same as long as the time interval is the same, ANY TIME INTERVAL. By the way, this is a joint distribution. So let’s write this distribution, If you’re modeling returns, you get $$ R_t = \\log S_t $$ and $$ R_{t+\\Delta t} - R_t = (\\mu - \\frac{\\sigma^2}{2}) \\Delta t + \\sigma \\Delta W_t $$ And $$ \\log S_{t+\\Delta t} = \\log S_t + (\\mu - \\frac{\\sigma^2}{2}) \\Delta t + \\sigma \\Delta W_t $$ The only randomness comes from Brownian motion, which has N(0, Δt) I know the distribution of this thing, so $$ \\log S_{t + \\Delta t} | S_ t \\sim N(\\log S_t + (\\mu - \\frac{\\sigma^2}{2} \\Delta t, \\sigma^2 \\Delta t) $$ This is a constant that I add to it, so the only thing that happens is the mean changes when I add to the normal. Now that I know this, I can write down the density. obviously, the logarithm of this thing is normal. Which means S is e to this thing. It’s just easier to write in the context of the logarithm $$ f(\\log S_{t+\\Delta t} | \\log S_t, \\theta) = \\frac{1}{2\\pi \\sigma^2 \\Delta t} e^{-\\frac{X - \\log S_t - (\\mu - \\tfrac{\\sigma^2}{2})\\Delta t)^2}{2\\sigma^2 \\Delta t}} $$ This is just the PDE of the normal And notice that there’s no t, just Δt. So the only thing that matters is S\_t and Δt. That’s what homogenous means. What we know is: ## The Feller Process This is any Ito process, any SDE, which is written like $$ dX_t = f(X_t, \\theta) dt + g(X_t, \\theta) dW_t $$ Any process where’s the coefficients here don’t depend on time, only the stochastic process and the parameter. The big deal is, any Feller process is homogenous. It turns out that as long as the coefficients dont’ depend on time, the solutions odn’t depend on time. I make all these parentheses because it’s kind of crucial for us. It’s not enough for us to have an equal distance between times, you need to have a homogenous process. If you have daily, then they’re different, 100 different functions for 100 different days. Now I just have one function, which I plug in at different points in time. The following is NON-homogeneous. $$ dX_t = f(X_t, t, \\theta) dt + g(X_t, t, \\theta) dW_t $$ I had a friend that was Russian and pronounce it gomogeneous. That normal density is the derivation as well. If we come back to it, it’s the same as $$ p(y, x, \\Delta t | \\theta) $$ This is the transition probability from \\(X = \\log S_t\\) to \\(Y = \\log S_{t+\\Delta t}\\) Notice I don’t care what Δt is, it could be a different one. $$ = \\frac{1}{\\sqrt{2\\pi \\sigma \\Delta t}} e^{-\\tfrac{(y - x - \\nu \\Delta t)^2}{\\text{not done}}} $$ This does work for GBM We can write down explicitly what it is. This function si solved, I write down the transition, I write the likelihood function, I take the logarithm. The joint distribution is a product. The logarithm is a sum of these functions, it’s just easier to deal with sums. In general, for other models than GBM, it is impossible to write an exact formula. So what do you do? A guy from Princeton in (1999, 2000\) named Ait Sahalia got famous for doing this. He came up with the idea of **Pseudo MLE**, also the **Approximate Likelihood Method**. There is a package that does this, which does not quote this. [https://www.princeton.edu/\~yacine/mle.pdf](https://www.princeton.edu/~yacine/mle.pdf) Reminder: there is a transition from X to Y in times Δt which depends on θ. We will replace p(y, x, Δt | θ) which we call pθ with a density hθ which depends maybe on some other parameters. We choose hθ in the following way: - has to be simple a constant is too simple, so it has to be related to this pθ, so: - hθ has to have the same moments as pθ This is a transition distribution in Y. The moments mean expected value of powers. We have this mgf. We know if two variables have same mgf, they have the same distribution. You can have two variables with several moments equal, but not the same distribution. This property only holds for ALL moments. So this wouldn’t be possible. I’m going to pick some hθ so we have *a few* moments the same, maybe 1, 2, 3\. The method in Ait Sahalia, we will discretize SDE using Euler method. We will approximate the integral with the LHS value times the increment, same as quadrature. We will cover Monte Carlo next week. $$ X_{t+\\Delta t} - X_t = \\int_t^{t + \\Delta t} f(X_t) dt + \\int_t^{t + \\Delta t} g(X_t) dW_t $$ We’re skipping θ for laziness, but exists in every function. This is the same as $$ = f(X_t) \\Delta t + g(X_t) \\Delta W_t $$ *if we approximate using quadrature*. We know the distribution of Brownian motion as N(0, Δt) Euler method says, that given X\_t, I know the distribution\! So $$ X_{t+\\Delta t} | X_t = X_t + f(X_t) \\Delta t + g(X_t) \\Delta W_t $$ And its distribution is $$ X_{t+\\Delta t} | X_t \\sim N(X_t + f(X_t) \\Delta t, g^2 (X_t) \\Delta t) $$ This is not the distribution, because we have an approximation on the integral with quadrature. Then we can approximate pθ as $$ h(\\Delta t, x, y) = \\frac{1}{\\sqrt{2\\pi g^2 (x) \\Delta t}} e^{-\\frac{-(y- x - f(x) \\Delta t)^2}{2g^2(x) \\Delta t}} $$ And if you give me f and g, and the parameters, I can technically write this down, very easily. It is simply multiplied like before. The x and y will be replaced by observations. How do you estimate analytically. You take the logarithm, which takes everything out and simplifies it. You get a bunch of the gs and the fs squared, but it’s just sums. If this would be linear, it would be simply applying an average. But nobody said this is linear. f and g can be complicated, so you would solve them with a nonlinear optimizer. ## Sim Diff Procedure First, let’s see two things. The O-U model: How do you solve it? Because it is solvable. Reminder: $$ dX_t = \\alpha(\\mu - X_t) dt + \\sigma dW_t $$ This is just notation for the stochastic integral, blah blah. The thing about this process, I can write it like $$ dX_t + \\alpha X_t dt = \\alpha \\mu dt + \\sigma dW_t $$ By the way, if there is no mean reverting here, it’s just called O-U. With mean reverting, it’s called mean reverting O-U. The alpha mu dt thing is gone in that case. This is the real trick, which comes from differentia equations. If you look at the LHS, you should remember something. You remember integrating factors, the way to solve first order diffeqs. And in fact that is a firs torder. What do you do with integrating factor? I don’t remember, I remember it existed, but it’s complicated, with some weird formula. But I do remember that if I take the derivative of \\(e^{\\alpha t}\\), then I’ll get \\(\\alpha e^{\\alpha t}\\) That’s a powerful feature, to make a constant appear out of nowhere. And we can also see another constant that appeared out of nowhere. So if we multiply everything here by that, you kind of get the derivative of the product to appear. $$ e^{\\alpha t} dX_t + X_t \\alpha e^{\\alpha t} dt = \\mu \\alpha e^{\\alpha t} dt + \\sigma e^{\\alpha t} dW_t $$ Now this is a stochastics process, so it’s a little different from the regular product rule. This process becomes $$ d(e^{\\alpha t} X_t) $$ I claim that. But how do you show this? What is the differential of two stochastic processes. That’s the Ito product rule. Which is the regular product rule plus the quadratic variation between the two terms. When you take the quadratic variation between stochastic process and a deterministic process, that ends up being zero. So actually we can get rid of the quadratic variation part, so it ends up being the same as the regular product rule. $$ = \\mu d(e^{\\alpha t}) + \\sigma e^{\\alpha t} dW_t $$ The next step is to integrate, from 0 to t. $$ e^{\\alpha t} X_t - X_0 = \\mu (e^{\\alpha t} - 1) + \\int_0^t \\sigma e^{\\alpha s} dW_s $$ Notice this is an explicit solution. $$ X_t = e^{-\\alpha t} X_0 + \\mu (1 - e^{-\\alpha t}) + \\int_0^t \\sigma e^{\\alpha (s-t)} dW_s $$ Just rearranging to get this here. We can’t really touch the dW stuff And this is random, because it has a stochastic integral. Another trick. Even for this simple expression, it is complicated to measure μ, α, or σ, because the stochastic integral is present. You can do the numerical thing and hope for the best. I learned this trick when I was in Romania. If you have this equation $$ X_t - X_0 = \\int_0^t \\alpha(\\mu - X_s) ds + \\int_0^t g(X_s) dW_s $$ I have some horrible expression like this, which I write as g, as complicated as you want. So what is the expected value of this? I get scared, because it’s so ugly. Because I have to know the pdf, and integrate x \* pdf, which is horrible. But actually, there’s a simple way to do this. It’s all based on the fact that stochastic integrals are martingales. Being martingales, they have the same expectation at any moment in time. And the process is equal to 0 at 0, so that whole thing ends up becoming 0\. The trick is to apply expectation everywhere. $$ \\mathbb{E}[X_t] - \\mathbb{E}[X_0] = \\mathbb{E}\\left[\\int_0^t \\alpha(\\mu - X_s) ds\\right] + 0 $$ The next step is to realize is that this is two integrals, which will commute as long as the thing inside is finite. And it has to be finite. $$ = \\int_0^t \\alpha (\\mu - \\mathbb{E}[X_s]) ds $$ And now this is an expectation, which is a number, it just depends on the time. We’ll define a function for this. $$ \\mathbb{E}[X_t] = u(t) $$ So we’re really solving this equation: $$ u(t) - u(0) = \\int_0^t \\alpha (\\mu - u(s)) ds $$ But this is now a deterministic equation which you can solve as diffeq using Calculus III. You take the derivatives, which makes it simpler, because it mimics what we just did. $$ du(t) = \\alpha(\\mu - u(t)) dt $$ This the same equation we did a moment ago $$ du(t) + \\alpha u(t) dt = \\alpha \\mu dt $$ And this is a first order diffeq, where we can use the same e thing. $$ d(e^{\\alpha t} u(t)) = \\alpha ue^{\\alpha t} dt $$ And now RHS integrates really easily, same as before $$ e^{\\alpha t} u(t) - u(0) = \\mu (e^{\\alpha t} - 1) $$ $$ u(t) = u(0) e^{-\\alpha t} + \\mu (1 - e^{-\\alpha t}) $$ And actually we didn’t need to do all of this, because on the previous slide, we had the solution. If we applied expectation directly to the Ito product rule we did, it would give the same formula, which is just more generalized. $$ \\mathbb{E}[X_t] = \\mathbb{E}[X_0]e^{-\\alpha t} + \\mu (1 - e^{-\\alpha t}) $$ And you can observe the long term behavior, that it goes to μ. That’s the theory. How do we estimate the stuff? ## Code I wrote this before coming to class, which is why I was late. This is part of a package called Sim.DiffProc, which you need to install. For some nonstandard packages, you need Rtools, looks like you don’t need it. Now they have a very nice Sim.DiffProc is not that good. Simulation of Diffusion Processes. They have a nice documentation for it. They can do more things than this. Here: ## What You Should Remember MLE Method Approximation of Joint distribution with little pieces Feller processes and what they are This is stuff that actually shows your advantage over students from competing programs. They don’t do anything like this. ## CDS and Credit Markets - URL: https://sharifhsn.dev/blog/cds-and-credit-markets/ - Structured data: https://sharifhsn.dev/api/posts/cds-and-credit-markets.json - Description: The FE-635 syllabus turns to credit as a traded asset through credit default swaps. A CDS exchanges a premium leg for protection against a defined credit event. The buyer pays the … - Date: 2025-03-31 - Exact published timestamp: 2025-03-31 - Topics: FX, CDS, Credit Markets - Categories: FX - Source: FE-635 \| Risk Engineering - Source URL: None The FE-635 syllabus turns to credit as a traded asset through credit default swaps. A CDS exchanges a premium leg for protection against a defined credit event. The buyer pays the spread while the reference entity survives and receives a loss payment after default, subject to the contract's recovery convention. The notes connect the CDS price to a survival curve and a recovery assumption. A quoted spread is therefore not a pure probability of default; it also contains funding, liquidity, and risk premia. The premium and protection legs must be discounted consistently and aligned on payment dates. This is why the risk-engineering workbook keeps instrument conventions explicit rather than hiding them behind a single “credit” input. ## The Cost of AI-Generated Images - URL: https://sharifhsn.dev/blog/linkedin-2025-03-28-the-cost-of-ai-generated-images/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-03-28-the-cost-of-ai-generated-images.json - Description: New AI images stun the world 🤯 but are they worth the cost? 🤔 - Date: 2025-03-28 - Exact published timestamp: 2025-03-28T13:00:08.066Z - Topics: AI, Creative Work - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7311371515738898434/ New AI images stun the world 🤯 but are they worth the cost? 🤔 AI images have flooded social media seemingly overnight, with the specific trend of “Ghiblization” (transforming images into the style of animation studio Studio Ghibli) becoming a popular meme. These images from ChatGPT 4o are much higher quality than previous AI images. That’s not a coincidence. 4o uses a different method of image generation than the old model, DALL-E. DALL-E was based around diffusion. This method models the noising of an image through a stochastic process, which has a closed-form sampling method. Then, it uses Anderson’s reverse diffusion to denoise. The score function is estimated by the UNet neural network which is trained on real images that are noised. The denoising process starts with pure Gaussian noise, then denoises at each step. It uses the attention mechanism just like LLMs do to recognize when a patch of noise looks like something in the prompt, and denoises it to fit the desired image. After only a few dozen steps, the image is denoised. This process has benefits and drawbacks. It’s very computationally efficient, as it renders the whole image at once in a parallelizable fashion. However, this independence between batches causes a lack of spatial reasoning, which is why diffusion images often have messed up hands or text. 4o uses autoregression instead. You can think of this as similar to what ChatGPT does, where each part of the image is generated based on previous parts of the image, left to right and top to bottom. This takes care of the spatial reasoning element, since each part of the image develops from earlier parts. But it requires a full forward pass through the transformer stack for each new image token, attending to all previous tokens. This makes the time complexity of the process O(n^3), with no opportunity for parallelization! This is why access to 4o’s image generation capabilities are being limited, and Sam Altman has said “our GPUs are melting”. Much hay has been made about the energy consumption of AI, and I generally think that these worries are overblown compared to other uses of energy. But autoregressive image generation seems to push that limit hard. What do you think? Are the higher quality images worth it? ## Hazard Rates and Credit Default Swaps - URL: https://sharifhsn.dev/blog/advanced-derivatives-week-09/ - Structured data: https://sharifhsn.dev/api/posts/advanced-derivatives-week-09.json - Description: Lambda hazard rate is convenient to work with. - Date: 2025-03-27 - Exact published timestamp: 2025-03-27 - Topics: Credit, Hazard Rates, Recovery Rates, Credit Default Swaps - Categories: Credit - Source: Advanced Derivatives - Source URL: None ## Hazard Rate Lambda hazard rate is convenient to work with. If we integrate over it, we can get the survival probability V(t) at time t. We can also get the probability of default Q(t) by doing 1 \- survival probability. ## Recovery Rate If we have default, we will lose all future coupons/payoffs. Typically we can recover something from the bond, depending on the structure of the debt. Bonds get paid before equity. The recovery rate is defined as the price of the bond immediately after default as a percentage of FV. Recovery rates DECREASE as default rates INCREASE. Recovery is not known in advance but it is estimated in a certain range based on historical data. We can get the implied probability of default for a bond using some qualities. - the bond price - CDS spreads (will discuss) CDS is basically an insurance instrument You can also get historical data to construct hazard rates Merton’s model also (will discuss) ## Bond Prices Obviously, there is a relationship between default and bond price. We can approximate default intensity over life of bond as $$\\frac{s}{1-R}$$ where s is spread of bond’s yield over risk-free rate and R is recovery rate. We can look at a little more exact calculation. Let’s say we have slide 13 from Part 1\. When we do pricing of a CDS, we will look more detail in this issue, where if defaults can happen at any time, we integrate over the domain. The goal is to calculate such probability of default. We have a coupon payment of $3. So the YTM is You can estimate risk-neutral default rate each year. It’s a bootstrap process. You start with lower maturity bonds, then get higher maturity bonds ## Credit Default Swaps Excess of n-bond yields of corporate bonds must equal CDS spread, otherwise there is arbitrage where you can either earn over the risk-free rate or borrow at less than the risk-free rate. The CDS bond basis is the spread \- the excess bond yield. From the arbitrage argument, the bond basis should be 0\. Historically, CDS bond basis \> 0\. One leg is the protection buyer, the premium leg, and the other is the When you value the premium leg, The other element is the survival probability. The second component is that in case of default, you have to pay the accrual ## Midterm Cheat Sheet - URL: https://sharifhsn.dev/blog/computational-methods-midterm-cheat-sheet/ - Structured data: https://sharifhsn.dev/api/posts/computational-methods-midterm-cheat-sheet.json - Description: Handwritten midterm reference for quadrature and Black-Scholes notation; several formulas are unclear in the source render. - Date: 2025-03-25 - Exact published timestamp: 2025-03-25 - Topics: Computational Methods, Midterm, Quadrature, Numerical Integration, Midpoint Rule, Simpson's Rule, Black-Scholes, Partial Differential Equations, Free Boundary - Categories: Computational Methods - Source: Handwritten CheatSheet.xopp; Drive-modified 2025-03-29. Approximate date within the Computational Methods in Quantitative Finance Docs history cluster (2025-01-21 through 2025-05-15). - Source URL: None The handwritten page is headed “Computational Methods in Quantitative Finance Midterm Cheat Sheet.” ## Quadrature / numerical integration Readable labels: “Quadrature,” “numerical integration,” “Midpoint Rule,” and “Simpson’s Rule.” The line beginning “for Taylor:” and the formulas under the rule labels are **unclear in the source**. ## Black-Scholes The page marks “PDE” and “Free Boundary.” The associated PDE, Greek symbols, and boundary equations are **unclear in the source**. The finite-difference grid and method derivations are covered in weeks 5 and 6. This supplement keeps the readable quadrature and section labels without repeating those derivations. ## GARCH Risk Statistics - URL: https://sharifhsn.dev/blog/garch-risk-statistics/ - Structured data: https://sharifhsn.dev/api/posts/garch-risk-statistics.json - Description: The risk-statistics portion of FE-635 introduces GARCH as a way to model time-varying volatility. Returns may have little serial correlation while their squared returns cluster: ca… - Date: 2025-03-24 - Exact published timestamp: 2025-03-24 - Topics: FX, GARCH, Volatility Estimation - Categories: FX - Source: FE-635 \| Risk Engineering - Source URL: None The risk-statistics portion of FE-635 introduces GARCH as a way to model time-varying volatility. Returns may have little serial correlation while their squared returns cluster: calm periods tend to be followed by calm periods, and shocks tend to persist. A simple GARCH(1,1) recurrence is $$\sigma_t^2=\omega+\alpha\epsilon_{t-1}^2+\beta\sigma_{t-1}^2.$$ The parameters describe long-run variance, reaction to a new shock, and persistence. Estimation and diagnostics matter as much as the recurrence; a fitted model should be checked against the horizon and the tail behavior of the risk report. ## LIBOR Market Model and Hazard Rates - URL: https://sharifhsn.dev/blog/advanced-derivatives-week-08/ - Structured data: https://sharifhsn.dev/api/posts/advanced-derivatives-week-08.json - Description: The LIBOR Market Model (or Brace-Gatarek-Musiela Model) - Date: 2025-03-20 - Exact published timestamp: 2025-03-20 - Topics: Fixed Income, LIBOR Market Model, Interest Rate Models, Hazard Rates - Categories: Fixed Income - Source: Advanced Derivatives - Source URL: None ## LMM The **LIBOR Market Model** (or **Brace-Gatarek-Musiela Model**) is a model constructed in terms of the forward rates underlying caplet prices. HJM was defined as a process for the forwards. Basically the benefit was that it could model the entire term structure, as long as you input all the volatilities for the term structure. But the problem is the calibration. Notation \\(t_k\\) kth reset date So t would progress like t\_0 \= 0, t\_1 \= 0.25, t\_2 \= 0.5… \\(F_k\\) forward rate between k and k \+ 1 m(t) index for next reset date at time t. So if t \= 0.4, then m(t) \= 0.5 δ\_k \= t\_k+1 \- t\_k So basically you can imagine that you have your forward curve, xi: ξ(t) \= volatility of F\_k(t) at time t Assume that we have only one factor This factor is the forward risk neutral process with respect to P(t, t\_k+1) Then, our process $$ dF_k(t) = \\xi_k (t) F_k(t) dW $$ The change is driven by the volatility, existing forward rate, and Brownian motion. Applying some change of numeraire, skipping some steps, you get the rolling forward rate process (in notes) Price of a ZCB at time t\_i which is discounted, the ratio of that to t\_i+1 $$ \\frac{P(t, t_i)}{P(t, t_{i+1})} = 1 + \\delta_i F_i(t) $$ If you apply the natural log, you get $$ \\ln P(t, t_i) - \\ln P(t, t_{i+1}) = \\ln [1 + \\delta_i F_i(t)] $$ This is the relationship between the ZCB and the forward. Then if you apply Ito’s lemma… and equate coefficient of dW something long… Then you can take the substitution of the forward process. $$ \\frac{dF_k(t)}{F_k(t)} = \\sum_{i= m(t)}^k \\frac{\\delta_i F_i(t) \\xi_i(t) \\xi_k(t)}{1+\\delta_i F_i(t)} dt + \\xi_k(t) dW $$ This is the process followed by the forward rate between t\_k and t\_k+1, in a risk-neutral world. In the limiting case, when this interval becomes smaller, this converges to HJM The idea is being able to calibrate this model. It can be simplified with: ## Simplified Model Assume that ξ\_k(t) function only on the number of whole accrued periods between the next date and t\_k. Define Λ\_i as the value of ξ\_k(t) (volatility of forward) when there are i such accrued periods. Then the volatility of the forward can be redefined as \\(\\xi_k(t) = \\Lambda_{k-m(t)}\\) which is a step function And such values like Λ\_i can be estimated from the volatilities used to value caplets in Black’s model. Recall that to value a caplet, we have $$ L\\delta_k P(0, t_{k+1}) [F_k N(d_1) - R_k N(d_2)] $$ blah blah blah d\_1 and d\_2 If we equate the variances between Black’s model and this, we get $$ \\sigma_k^2 t_k = \\sum_{i=1}^k \\Lambda_{k-i}^2 \\delta_{i-1} $$ Example in the slides for converting Black volatilities to Lambda, use to check ## Implementation How would I use this model to price a bond. If you observe the volatility, then you can price a bond under this? This would be done via Monte Carlo simulation $$ \\frac{dF_k(t)}{F_k(t)} = \\sum_{i=m(t)}^k \\frac{\\delta_i F_i(t) \\Lambda_{i - m(t)}\\Lambda_{k-m(t)}}{1 + \\delta_i F_i(t)} dt + \\Lambda_{k-m(t)} dW $$ Then, by Ito’s lemma, $$ d\\ln F_k(t) = \\left[\\sum_{i=m(t)}^k \\frac{\\delta_i F_i(t) \\Lambda_{i-m(t)} \\Lambda_{k-m(t)}}{1 + \\delta_i F_i(t)}\\right]dt + \\text{same thing as above} $$ Then we can approximate the drift by approximating \\(F_i(t)\\) and \\(t\\) Then, $$ F_k(t_{j+1}) = F_k(t) \\exp\\left[\\biggl(\\sum_{i=j+1}^k \\frac{\\delta_i F_i(t_j) \\Lambda_{i-j-1}}{1 + \\delta_i F_i (t_j)} - \\frac{\\Lambda^2_{k-j-1}}{2}\\biggr) \\delta_j + \\Lambda_{k-j-1} \\epsilon \\sqrt{d_j}\\right] $$ Slide 8 has more calculations that I can use for testing. ipynb file is provided Each approximation of the drift the forward rate remains constant. If we define the rolling forward in the risk neutral world, it allows us to discount bond from one date to the next one. In terms of simulation, what’s happening in the code is that if you assume that we want to simulate a zero curve with N accrued periods. On each trial, we start with the forward rate at time 0, which is calculated from the initial zero curve. F\_0(0), F\_1(0), … F\_N-1(0) Then we can use an approximation formula (described above) to calculate F\_1(t\_1), F\_2(t\_1)... Limitations of the theoretical model ## Credit Risk LMM is a little complicated, this is just an introduction to the topic. Credit risk is the probability that a debtor will default on their debts. We have always assumption that the cash flow is 100% likely i.e. with treasury bonds. For single-name derivatives, we will look at CDS and CDOs. Then we will look at portfolios. We are interested in expected loss, with a discrete number of names in the portfolio, to estimate the losses in order to price the portfolio correctly. Rating agencies will rate the credit risk of bonds. S\&P says AAA, AA, A, BBB, BB, B, CCC, CC, C Moody’s has Aaa, Aa, A, Baa, Ba, B, Caa, Ca and C Bonds with ratings of BBB and above are investment grade These are measured as bands of probabilities. ### Question about Diversification By diversifying your portfolio into multiple assets, you decrease the credit risk. Then you can securitize the portfolio by splitting it into tranches, and then get the CDO. And then payment in the bonds You can have nonlinear dependence between variables, and tail dependence between variables which is conditional. Then the probability of default, what is the likelihood of this house being on fire if this other house is on fire? If the answer is not 0, then you have tail dependence. We will look at different models that measure this. ### Hazard Rates Also known as default intensity, this is the probability of default for a certain time period conditional on no earlier default. Unconditional default probability is from 0\. We are given a table that has cumulative default rates. In order to calculate the probability that A defaults in the first year, you take the actual probability. Then if it’s the second year, we’re concerned with the cdf(2) \- cdf(1). Then the probability of survival is 1 \- this This is all unconditional default probability. Then the conditional probabilities are known as default intensities or hazard rates The unconditinal probability of default within a certain default vs the survival probability. The survival probability is V(t) $$\\lambda(t) \\Delta t \= \\frac{V(t) \- V(t \+ \\Delta t)}{V(t)}$$ Then you get the ODE $$\\frac{dV(t)}{dt} \= \-\\lambda(t) V(t)$$ $$V(t) \= e^{-\\int\_0^t \\lambda(t) dt}$$ Then the survival probability can be found by integrating over the hazard rate. So if you have a specification of the hazard rate, you can find the cumulative survival probability. Then we can also discuss the cumulative default probability. We can use the CDS to determine the piecewise constant hazard rates and then construct the survival probabilities. So this is a calibration in which we can price the risk of default through CDS. First we will look at bonds. ## Stochastic Volatility, Jumps, and Fourier Transforms - URL: https://sharifhsn.dev/blog/computational-methods-week-09/ - Structured data: https://sharifhsn.dev/api/posts/computational-methods-week-09.json - Description: Theoretical transformations, very technical - Date: 2025-03-18 - Exact published timestamp: 2025-03-18 - Topics: Computational Methods, Stochastic Volatility, Jump Processes, Poisson Process, Fourier Transform - Categories: Computational Methods - Source: Computational Methods in Quantitative Finance - Source URL: None Theoretical transformations, very technical ## Stochastic Volatility Models How do you solve a Heston model using Fourier transforms? ### Hull-White The very first stochastic volatility model introduced is the Hull-White model. $$dS\_t \= rS\_t dt \+ \\sqrt{y\_t} S\_t dW\_t$$ $$dY\_t \= \\mu\_Y Y\_t dt \+ \\sigma\_Y Y\_t dZ\_t$$ Has no solution:( If the Brownian motions are uncorrelated, then there is no leverage effect, because that effect is a correlation between volatility and returns. But just because the Brownian motions are uncorrelated, doesn’t mean the processes are uncorrelated, they incorporate each other. ### Leverage Effect The leverage effect is the perceived correlation between returns and volatility, and news and returns. The idea is that positive news causes the stock to go up, which makes return go up, and vice versa for negative news. However, the two effects are not the same. If there is good news the stock goes up, if there’s bad news, the stock goes down much more proportionally. This is called the **leverage effect**, which is a negative correlation between volatility and returns. Basically, the news increases vol, and as vol increases, returns go down. You need to have stochastic volatility to have correlation between stochastic and deterministic processes. I’ll mention two more stochastic volatility models. ### SABR We discussed SABR earlier, so this is a reminder. It looks like $$dS\_t \= rS\_t dt \+ \\sigma\_t S\_t^\\beta dW\_t$$ This is the most general form of the process, most that you see in practice have r \= 0 because they are created by physicists who don’t like complicated models. The sigma is the stochastic process, which has the process $$d\\sigma\_t \= \\alpha \\sigma\_t dZ\_t$$ The W and Z are two Brownian motions which can be correlated with rho. This is also called the stochastic alpha beta rho model (SABR) for the three parameters and the stochastic volatility. The authors of this made it so that once you calculate sigma, you can plug it into Black-Scholes and reuse all your old code. ### Constant elasticity of variance (CEV) Pioneered by Peter Carr (friend of the show) and Madan. $$dS\_t \= rS\_t dt \+ \\sigma S\_t^{\\frac{\\beta}{2}} dW\_t$$ If you think about this in terms of stochastic models. This beta is strictly less than 2, and sigma is greater than 0\. You ca write this as $$\\sigma S\_t S\_t^{\\frac{\\beta \- 2}{2}}$$ Which means that the last term is the actual stochastic volatility. That’s literally driven by the price evolution. This is inversely proportional to the value of the stock. If you Ito this with log S\_t, you get $$dX\_t \= \\left(r-\\frac{\\sigma^2 S\_t^{\\frac{\\beta-2}{2}}}{2}\\right) dt \+ \\sigma S\_t^{\\frac{\\beta-2}{2}} dW\_t$$ Then the second term is the vol. So the variance should be the vol squared. $$\\mathbb{V} \= \\sigma^2 S\_t^{\\beta \- 2}$$ If you compute the derivative the change in the variance with respect to stock. $$\\frac{\\partial \\mathbb{V}}{\\partial S} \= \\sigma^2 (\\beta \- 2\) S\_t^{\\beta \- 3}$$ You can rewrite this as $$\\sigma^2 S\_t^{\\beta \- 2} \\frac{\\beta \- 2}{S\_t}$$ The point of doing this is to see that these left two terms are the exact same as the variance. $$\\frac{\\partial \\mathbb{V}}{\\partial S} \= \\mathbb{V} \\frac{\\beta \- 2}{S\_t}$$ Which is the same thing as saying $$\\frac{\\partial \\mathbb{V}}{\\mathbb{V}} \= (\\beta \- 2)\\frac{\\partial S}{S}$$ The change in variance is proportional to the change of stock. The variance move elastically, proportional to the way the stock moves. And they change inversely, because beta is less than 2\. And actually, if beta \= 2, then you get GBM in the formula. $$dS\_t \= rS\_t dt \+ \\sigma S\_t dW\_t$$ Which makes sense because that assumes constant volatility, where $$\\frac{\\partial \\mathbb{V}}{\\mathbb{V}} \= 0$$ When beta \= 1, you get CIR $$dS\_t \= rS\_t dt \+ \\sigma \\sqrt{S\_t} dW\_t$$ Changes in the stock become actually 1:1 with the stock in this case. ## Jump Processes Lonon is an expert on jump processes. ## Poisson Process Basically, $$N \\in \\{0, 1, \\ldots \\}$$ where at time t, $$N\_t \\sim Poisson(\\lambda t)$$ Lambda quantifies the expected number of values for when t \= 1\. That’s how you scale it. The question is, how do you simulate this? There are two ways to do this: Probability and Stochastic Processes (Florescu) has a lot of information on this. I suggest you pirate my book because the publishers are thieves. Method 1: If $$X\_1, X\_2, X\_3, \\ldots X\_n$$ are iid Exp(1/lambda) Basically the expected amount over lifetime is 1/lambda. Then we let $$T\_1 \= X\_1, T\_2 \= X\_1 \+ X\_2, T\_3 \= X\_1 \+ X\_2 \+ X\_3, \\ldots$$ What I’m doing here is defining the event times for the Poisson process. At some point t, if you were to plot the process, then you would get a bunch of jumps up to t. Then $$N\_t \= \\max\_n \\{T\_n \\leq t\\}$$ The only question is if it’s included or not, so le’ts be careful. Actually, it’s $$N\_t \= \\inf\_n \\{T\_n \> t\\}$$ You can understand how to create this\! You simply generate these exponentials from the distributions, and you know the times from the sum. You give me the process, and the time t. The jump is of value 1, so it always jumps by 1\. There is a marked Poisson process or compound Poisson process, where instead of moving it by 1, you move it by a random variable, and then you sum those. Once you understand this it’s very simple, you just generate a series of random variables and assign them to each T. How is this useful? In two lectures, we will learn about Monte Carlo simulations. These typically don’t have jumps. But if you add jumps to them, you get a jump process. You will get your T\_1, T\_2, and T\_3 etc. You will be going up and down by random quantities at those times. These jumps: Trump announces tariffs at each time T, and it goes up and down based on what country he tariffs. There is some process which is not Jumpy, looks more like a Brownian motion. To introduce the jumps, you simply shift the value by the value of the random variable. The exponential distribution starts at 1/lambda at t \= 0 and then goes down. for Exp(1/lambda) Most of the times, you get a small value. Secret: if you play video games, and play whatever discrete events that happen in time. If you play a gacha game and put coins in there to get the good Pokemon. You keep getting a crappy Pokemon, and suddenly they give you a good one. But now you say it stops because you have it. The time between Pokemons is a random variable. If you generate the Poisson process for this, and if you look at the path, and you think about it logically, if it happens 5 per hour, the jumps should happen evenly. But this is not how it really looks. If you look at the pdf, you’re much more likely to have smaller intervals than larger intervals. Method 2 is based on the following two results. If we look at interval \[0, t), the number of events is distributed as Poisson(lambda t) Let’s say you have events that happen once a day. Trump issues things once per day. (I’m a Republican so don’t pick on me) If you want to generate how many things this guy says this week, you make a random variable with parameter t \= 7, lambda \= 1\. This Poisson random variable. Given there are N events in the interval \[0, t) that I’m generating, the times of the events (and this is proven in the book), are the order statistic from N uniform \[0, t\] random variables. That sounds fancy, but basically it says… An **order statistic**. If you have n random variables iid, and you take them X\_1, X\_2, X\_n, the order statistics ordered like $$X\_{(1)} \\leq X\_{(2)} \\leq … X\_{(n)}$$ and this is just an ordered list. So the first order statistic is the smallest number. The order statistics have a distribution which depends on the original distribution, and it’s actually quite simple to work with them. Then generate random variable $$N \\sim Poisson(\\lambda t)$$ Then generate random $$N \\sim Uniform\[0, t\]$$ Let’s say Poisosn happesnt ob e 10\. I generate a variable Generate the uniform, look at the numbers, then list them smallest to largest. Let’s say it’s 1.1. It took Trump 1.1 days to say the first stupid things. Then you have 3, so it took him another day to say something else. These are the times, then the magnitudes come from them. rpois(1, 7\) for example, gives you 6 N=rpois(1,7) Now you generate seven uniforms runif(N, 0, 7\) This gives you the times, which are unsorted. Now you sort them sort(runif(N,0, 7)) Poisson gives us the number of events, and uniform gives us the actual times of the events. ## Transformation Methods ### Laplace Transform This is very familiar to probabilists. This is also called the moment generating function. If you define X as a random variable with pdf f(x) Then we define the mgf $$M\_X: \\mathbb{R} \\rightarrow \[0, \\infty)$$ $$M\_X(t) \= \\mathbb{E}\[e^{tX}\]$$ If X has a pdf, then this is also defined as $$= \\int\_{-\\infty}^\\infty e^{tx} f(x) dx$$ This was invented by Laplace, but he invented it in physics, for functions that were positive support, from 0 to infinity. So it’s a little different. The mgf is called so because if you take the derivative of the function with respect to t, you get the moments. (You need to prove that the derivative commutes with the integral, not that difficult). $$M’(t) \= \\frac{dM\_X(t)}{dt} \= \\int\_{-\\infty}^\\infty xe^{tx} f(x) dx \= \\mathbb{E}\[Xe^{tx}\]$$ And then $$M’(0) \= \\mathbb{E}\[X\]$$ The P\&SP book will cover this in more detail. The Laplace transform of f: (0, infinity) to R, looks like $$Lf(t) \= \\hat{f}(t) \= \\int\_0^\\infty e^{tx} f(x) dx$$ The only difference is that it’s positive. You can always write mgf as sum of two Laplace transforms. The Laplace transform has this nice **inversion theorem**. Moving past all the details, you should remember it as: If f has Laplace transform Lf, then $$f(x) \= \\frac{1}{2\\pi i} \\lim\_{T \\rightarrow \\infty} \\int\_{C \- iT}^{C \+ iT} e^{tx} Lf(t) dt$$ This limit exists only for certain Lf(t), and even if it exists it’s ugly. So undergrads will look at the table of Laplace transforms. [Table of Laplace Transforms](https://web.stanford.edu/~boyd/ee102/laplace-table.pdf) Stanford uses t and s, Florescu uses x and t. This is useful because if you take the derivative of the function, and apply the Laplace transform, it becomes a polynomial. It becomes tF(t) \- f(0) If you take the nth derivative, you get a bunch of derivatives evaluated at 0\. It makes it easier to solve. It’s very useful for solving diffeqs. In practice, this thing has two problems. The specific doesn’t exist. With small exceptions, the equation is too complicated to get the value of the function that corresponds to it. If you’re interested in applyin this and want to work with Laplace transform, use Mathematica which Dragos buys for Stevens ## Fourier Transform There is an equivalent to this in probability, which is the **characteristic function**, which is more complicated than the transform. What is the Fourier transform? For f(x): $$f: \[0, \\infty) \\rightarrow \\mathbb{R}$$ $$F(t) \= \\int\_{-\\infty, \\infty} e^{-itx} f(x) dx$$ The only difference is the introduction of the i thing. In general, you can apply Euler’s identity for $$e^{ia} \= \\cos a \+ i \\sin a$$ So this becomes $$F(t) \= \\int\_{-\\infty, \\infty} \\cos(tx) f(x) dx \- i \\int\_{-\\infty, \\infty} \\sin(tx) f(x) dx$$ You can calculate this using real integrals. Laplace is actually harder than this. What is the connection with the characteristic function? $$\\varphi\_x(t) \= \\mathbb{E}\[e^{itX}\] \= \\int\_{-\\infty}^\\infty e^{itx} f(x) dx$$ The only difference is that the characteristic doesn’t have a minus, which is not a big deal. The minus doesn’t mean anything, really. What is the connection between the characteristic and the moment? We can calculate moments from characteristic function. $$\\varphi\_x’(t)\\frac{d}{dt} \\varphi\_X(t)$$ You have to prove this works, and then take the complex function, which is not that hard. $$= \\mathbb{E}\[(iX) e^{itX}\]$$ It’s kind of like you go inside and take the derivative as normal. But if you take this $$\\varphi\_x’(0) \= i\\mathbb{E}\[X\]$$ This becomes slightly more complicated, because you get the powers $$\\varphi\_x’’(0) \= i^2 \\mathbb{E}\[X^2\]$$ And this continues in general. What is the advantage, why do we do this Fourier transform and not stick to the Laplace transform? This Fourier transform always exists. And there is also an inverse Fourier transform. It is basically the same idea. You get a diffeq and a simpler equation, and then the inverse gives you the solution. There is something more that exists here\! That is **discrete Fourier transform**. That is the big deal. Generally speaking, you will run into the same problem. You get this horrible expression from the Fourier transform, and you can’t get the pdf from it. Engineers have invented this approximation. You express the original function in terms of cosines and sines, and then you know the transforms and inverse transforms from there. [Table of Fourier Transform Pairs](https://engineering.purdue.edu/~mikedz/ee301/FourierTransformTable.pdf) I had an argument a long time ago… The whole point is, when you do a Fourier transform, you go into the frequency domain of your function. If your function is a sinus, you get one value. If you have a combination of sinuses, then you get a multitude of frequencies. The whole point of the DFT is that you express the function through the bases of sinuses. [Discrete Fourier transform](https://en.wikipedia.org/wiki/Discrete_Fourier_transform) There’s math here, but nobody actually uses it. What we do instead, is that we call a package and say this is my function, calculate the DFT, put it into an equation, solve it, calculate IFT, and then just do it that way. ## Applications of Fourier Transform Peter Carr was a guy at Bloomberg, and him and Ionut argued. He said he was solving stochastic vol formula analytically, Ionut says this is not possible. Carr admits that it’s just very fast. This is used to solve the Heston model. $$dS\_t \= S\_t(r dt \+ \\sqrt{V\_t} dW\_t)$$ $$dV\_t \= K(\\theta \- V\_t) dt \+ \\sigma \\sqrt{V\_t} dZ\_t$$ W and Z can be correlated with rho. Original Heston model is uncorrelated, there is an extension by Wiggins which is correlated. The solution is not that complicated. There is also something called the Feller condition. In order for this Heston model to be nicely behaved (homogenous), you should have $$2K\\theta \> \\sigma^2$$ The problem is, given this model for my stochastic process, which is actually quite realistic, we want to find the price of an option. t is time now S is stock price K is strike price T is time of maturity r is risk free rate NEW: V is the value of this variance process $$C(t, S, K, T \- t, r, V)$$ If you consider the other points observable, then this is a function of C(t, S, V) We used to solve things as $$ t \\in \[0, T\]$$, $$S \\in (0, infty)$$, $$\\mathbb{V} \\in (0, \\infty)$$ Now we solve this equation with respect to three parameters. Call option \= $$\\mathbb{E}^Q \[e^{-r(T \- t)} (S\_T \- K)\_+ | \\mathcal{F}\_t\]$$ This is the general formula. You can write this as, by splitting into two different situations, when in the money and out of the money $$=\\mathbb{E}^Q\[e^{-r(T-t)}(S\_T \- K)\_+ \\mathbb{I}\_{\\{S\_T \> K\\}}$$ Where we disregard the out of the money part because it’s worthless. $$= \\mathbb{E}^Q \[e^{-r(T-t)} S\_T \\mathbb{I}\_{\\{S\_T \- K\\}}|\\mathcal{F}\_t\] \- K\\mathbb{E}^Q\[e^{-r(T-t)} \\mathbb{I}\_{\\{S\_T \> K\\}}\]$$ Then we can take out the e term because it’s just a number, $$= e^{-r(T-t)} \\mathbb{E}\[S\_T \\mathbb{I}\_{\\{S\_T \- K\\}} | \\mathcal{F}\_t\] \+ Ke^{-r(T-t)}\\ldots$$ Then this first term will be considered P1(t, S, V), and the second term is P2(t, S, V). What he said is if you notice the very first property of the Fourier transform is that it is linear. So both P1 and P2 must solve the Heston PDE. Then let’s apply the Fourier transform to the original Heston PDE. This is more complicated than just going from t to x as before, because we have three variables. So we do: $$\\hat{f}(\\phi, x, v)$$ So he postulated that $$\\hat{f}(\\phi, x, v) \= e^{C(\\tau, \\phi) \+ D(\\tau, \\phi)v+i\\phi x}$$ This has some theoretical reasoning, but he says that once you plug it in, there is a solution. ## Next Week Estimating parameters, optimizations. ## Short-Rate Calibration and HJM - URL: https://sharifhsn.dev/blog/advanced-derivatives-week-07/ - Structured data: https://sharifhsn.dev/api/posts/advanced-derivatives-week-07.json - Description: We’ll look specifically at Vasicek and CIR. Then we will make some remarks on the Hull-White two factor model, then some more complex model. - Date: 2025-03-13 - Exact published timestamp: 2025-03-13 - Topics: Fixed Income, Calibration, Vasicek, Cox-Ingersoll-Ross, Hull-White, HJM - Categories: Fixed Income - Source: Advanced Derivatives - Source URL: None ## Calibration We’ll look specifically at Vasicek and CIR. Then we will make some remarks on the Hull-White two factor model, then some more complex model. The code for this in a package in R. ### Vasicek First we will simulate, then look at parameters. $$ dr_t = a(b - r_t) dt + \\sigma dW_t $$ We will use the Euler discretization scheme. $$ r_{t + \\Delta t} - r_t = a(b - r_t) \\Delta + \\epsilon_{t + \\Delta} $$ Then we will make a Binomial tree that uses such discretization. where $$ \\epsilon_{t+\\Delta} \\sim N(0, \\sigma^2 \\Delta) $$ Then in general, $$ r_{i\\Delta} = r_{(i - 1)\\Delta} + a(b - r_t) \\Delta + \\epsilon_{t + \\Delta} $$ ### Cox-Ingersoll-Ross $$ dr_t = a(b - r_t) dt + \\sigma \\sqrt{r_t} dW_t $$ Then do the same thing. ### Two-Factor Then the two factor Vasicek model can be defined with correlated factors. There’s a short rate and long rate factor. One such method of calibration is the **maximum likelihood estimator** (MLE) For Vasicek with real world data, consider sample short rates \\(r_0, r_{\\Delta}, r_{2\\Delta}\\), for Vasicek \\(f(r_{t+s} |r_t)\\) normal density with \\(\\mathbb{E}[r_{t+s}|r_t] = b + (r_t - b)e^{-as}\\) and $$ \\mathbb{V}[r_{t+s}|r_t] = \\frac{\\sigma^2}{2a} (1 - e^{-2as}) $$ Then the MLE method determines this for the distribution. It’s a function of parameters given the Typically you take the ln of the likelihood and determine the maximum given the function. $$ \\alpha^{*} = (1 - e^{-2a\\Delta}) b $$ $$ \\beta^{*} = e^{-a\\Delta} $$ $$ \\sigma^{*} = \\sqrt{\\tfrac{\\sigma^2}{2a} (1 - e^{-2a\\Delta})} $$ $$ \\mathbb{E}[r_{i\\Delta} \\mid r_{(i-1)\\Delta}] = \\alpha^{*} + \\beta^{*} r_{(i-1)\\Delta} $$ $$ \\mathbb{V}[r_{i\\Delta} \\mid r_{(i-1)\\Delta}] = (\\sigma^{*})^2 $$ Then we take the natural log of the likelihood for the normal density plugging in these parameters. MLE estimates $$ a = -\\frac{\\ln(\\hat{\\beta}^{*})}{\\Delta} $$ $$ b = \\frac{\\hat{a}^{*}}{1 - \\hat{\\beta}^{*}} $$ ## Risk-Neutral Calibration Assume that we have ZCB prices/discount factors (same thing). The price of a ZCB under the Vasicek model is $$ P(t, T) = e^{A(t, T) - B(t, T) r} $$ where $$ B(t, T) = \\frac{1}{a^{*}} (1 - e^{-a^{*}(T-t)}) $$ $$ A(t, T) = (B(t, T) - (T - t)) (b^{*} - \\frac{\\sigma^2}{2(a^{*})^2}) \\ldots $$ for the calibration, you would want to minimize for n bonds $$ \\sum_{i=1}^n (P_{\\text{Vasicek}} - P_{\\text{market}})^2 $$ You might use a nonlinear optimizer. ## Hull White Two Factor $$ dr = [\\theta(t) + u - ar] dt + \\sigma_1 dW_1 $$ $$ du = -bu + \\sigma^2 dW_2 $$ where we have another process u, and two volatilities/brownian motions. u in this context is a random mean reversion level. The previous model fits the term structure at time 0 and defines such mean reverting level, which is subject to some randomness. We have some historical data for htis. You also have ρ the correlation between dW\_1 and dW\_2, and the rest are constants. Spot rate volatility structure for Hull-White two-factor model. $$ \\sigma_R(t, T) = \\frac{1}{T-t} \\sqrt{[B(t, T)^2 \\sigma_1]^2 + [C(t, T)\\sigma_2]^2 + w\\rho\\sigma_1\\sigma_2B(t, T)C(t, T)} $$ Here is an example: Price a 5 year zero coupon bond after one year i.e. four years remaining to maturity. It has the following parameters: t \= 1, T \= 5, r(1) \= 0.05, u(1) \= 0.01. In order to calculate, all the implementation is in the PDF. First calculate B(1, 5\) B(0, 5\) B(0, 1\) C(1, 5\) C(0, 5\) C(0, 1\) ln A(1, 5\) P(1, 5\) Then you can construct the term structure will fit at time 0, and the evolution of this structure. Then in this case, we have an analytical solution for σ \= 0.0110 10 year and 15 year interest rates will have a lot of correlations, you can graph correlations against interest rate times and get the term structure. ## Limitations of One-Factor or Two-Factor Models Most involve only one factor of uncertainty. The models do not give the freedom in choosing the volatility structure. Looking at a simple example of rates, we can see that we have different volatilities for maturities, and these tend to be correlated. Based on these considerations, we need a more general approach in specifying the volatility environment and allow multiple factors This is the HJM Model (Heath Jarrow, and Marton Model) In the previous short rate models, we had a single source of randomness, Brownian motion, and the instantaneous correlation between forward rates and maturities is equal to 1\. We will consider some notation. P(t, T) is the price of ZCB of $1 at time t maturing at T. Ω will be the vector of past and present values of interest rates and bond prices at time t that are relevant for determining bond price volatility at that time, basically describes the information set. Then ν(t, T, Ω) this is nu, the volatility of P(t, T) The process for P(t, T) looks like $$ dP(t, T) = r(t) P(t, T) dt + \\nu (t, T, \\Omega_t) P(t, T) dW(t) $$ Then we have some characteristics at maturity. The forward rate is determined by the derivative with respect to the ln change in ZCB over time. The dynamics of the forward curve (ZCB) are only dependent on the volatility. Then we can say this is a risk-neutral process for the forward f that depends only on ν. Then we have a process for the instantaneous forward dF which is only determined by the volatility structure. So if we have ν, then we have the term structure There is a link between the dirft and the standard deviation of the instantaneous forward rate. But the extension to several factors with the market is expressed int erms of the instantaneous forward rate, which is not directly observable, and it’s difficult to calibrate. And also the process for the short rate in HJM model is non-Markovian. What is the connection between the HJM model and the previous short rate models? We look just at the single factor model. If you integrate the instantaneous forward rate, this is the actual short rate. And generally you can see how the terms depend on each part of the ν expression. That’s why r is non-Markovian, a non-recombining tree. ## American Options and Cubic Splines - URL: https://sharifhsn.dev/blog/computational-methods-week-07/ - Structured data: https://sharifhsn.dev/api/posts/computational-methods-week-07.json - Description: Cubic Splines, you solve something that is like a tridiagonal system - Date: 2025-03-11 - Exact published timestamp: 2025-03-11 - Topics: Computational Methods, American Options, Linear Complementarity, Successive Over-Relaxation, Cubic Splines - Categories: Computational Methods - Source: Computational Methods in Quantitative Finance - Source URL: None ## Three New Topics Cubic Splines, you solve something that is like a tridiagonal system Artificial Neural Network multilayer perceptron… Copulas (useful for 680\) ## American Option Valuation page 192 of the QF book the whole point of this is that you can exercise this at any time. So when do we exercise, at time τ? $$\\tau \\in \[0, T\]$$ There is a short, bad theorem that will help us. If the underlying process does not pay dividends, and is continuous, then the price of an American call and a European call are identical. You can prove this (the proof in the book is incorrect): Clearly, $$C\_A(S, t) \\geq C\_E(S, t)$$ American call for price S and time t, and European, because the American has extra on top of the call. And this is also true for the Put. $$C\_A(S, t) \\geq (S\_t \- K)\_+$$ $$P\_A(S, t) \\geq (K \- S\_t)\_+$$ You can argue this from a no-arbitrage argument, because if this were not true, you could instantly buy the option and exercise it and make money. In general, no-arbitrage arguments go by assuming that there is a strict inequality, then showing how you can buy low and sell high and make money instantly. Furthermore, $$C\_A(S, t) \\geq S\_t \- Ke^{-r(T-t)}$$ Why? We can do another non-arbitrage argument. You can put some money in a bank and get that K term risk-free. If we take the opposite of this, we can buy the cheap thing and sell the expensive thing. You buy the call option and the K ZCB (which becomes negative when you move it to the other side). First you short sell 1 share of stock (S\_t) Then you borrow Kert and put it in a bank Then you buy 1 call Because of this inequality, I will receive S\_t, and take my portion of the money I get and after taking away the call premium and the Kert, I’m left with a positive quantity. Everything with a minus is a liability, everything with a plus I have. At time T, my S becomes S\_T, which I have to give back because I short sold it. Then I will receive K and it will make up the balance Note that Kert is less than K, because e is raised to a negative exponent. And therefore $$C\_A(S, t) \\geq S\_t \- Ke^{-r(T \- t)} \\geq S\_t \- K$$ Because the value is always greater than the early payoff, then you should never exercise it early. The derivation in the book is wrong. This doesn’t work for puts. This result will also hold for continuous dividends, not discrete dividends. ### Free Boundary To understand the American option problem, we have to understand the free boundary problem. In the differential equation described, you have a system of equations you have to tie down with boundaries. There are two types of boundaries, Neumann in terms of the actual function, Dirichlet which is expressed in the derivative. This is one boundary on the top and one boundary at the bottom (and the terminal condition). These are tied down to these curves. Now a free boundary problem (which the American problem is) has another condition. This condition is not in a specific location. When are you going to exercise? When the expected future value of my option is going to be equal to or less than the value when I exercise now. If I exercise now and make more money, I should do that. But you don’t know what value of S and t will give you this. The problem is that you have to find this boundary. In the book, it refers to the physics problem. A lot of the applications of these math problems come from physics. In our case, let’s say I price an American Put. The payoff is at $$(K \- S\_T)\_+$$ You can prove that because the form of the function is monotonic. This is the final payoff, so it decreases as S increases, and then stays the same, so it is not. Then the value of the put option is also monotonic. There exists a price for the stock $$S\_f(t)$$ At any time t, this stock price exists, the “frontier” price (another name for boundary) we exercise if $$S\_t \< S\_f(t)$$ and we hold if $$S\_t \> S\_f(t)$$ What is the point of this? It’s a put option, so you make money if the stock price goes down. If the stock price is high, you make no money, and you want to hold. However, if the price is too low, maybe it will bounce up, so I should probably exercise. If you look instantaneously, like a fraction of a second right before maturity, and you should definitely exercise if the option is in the money. So there’s a region that has no exercise and a region where you do exercise. There’s a theorem: $$\\frac{\\partial P\_A}{\\partial S}(S\_f, t) \= \-1$$ The derivative right at the frontier is \-1. If you’re exercising at the frontier, the value you get is K \- S\_f, so you have take the derivative with respect to S\_f, it becomes \-1. There are a couple more steps in the actual proof because you need to say that it’s continuous and so on but this is the basic idea. If $$S\_t \\leq S\_f$$, then we exercise and get immediately $$(K \- S\_t)\_+$$ Then the free boundary problem tells us: $$P(S\_F, t) \= (K \- S\_f)\_+$$ and $$\\frac{\\partial P\_A}{\\partial S}(S\_f, t) \= \-1$$ As opposed to the other boundary problem, where it’s at fixed T, and you don’t know what S\_f is. ### Linear Complementarity Problem (LCP) $$(\\frac{\\partial P}{\\partial t} \+ \\frac{1}{2} \\sigma^2 S^2 \\frac{\\partial^2 P}{\\partial S^2} \+ rS \\frac{\\partial P}{\\partial S} \- rP) (P\_A(S, t) \- (K \- S)\_+) \= 0$$ For the first parentheses, this is geq 0, and the second parenthesis is also geq 0 . On each side of the space, one of them is 0\. It’s the same as the tree. You went to a point, and if that point is more worth it to exercise, you would store the value that comes from the tree. In American, you’re never going to solve the early exercise. Go back to your code for the American put and make the notes where you early exercise in red, and where you don’t exercise in black. In the lower part of the tree, you always exercise, and in the higher part you never do. Left paren is you cannot exercise, and right paren is where you do. If you don’t exercise, the fair value is the european value. If this stock value happens to be falling, then you exercise and right paren takes precedence. We set up the problem so that it takes over when you exercise. And then we also have the free boundary problems mentioned earlier. Basically, we solve the LCP using finite difference. In this, there are a bunch of steps that are reducing the problem, mentioned in the book like logarithm transformation, time transformation, and then you get equation 7.4.1 When you look at the finite difference method, you are looking backwards from the b\_i column at the end, and then u\_{i+1} (however we are going forward not backward) And then you solve it with AUi+1 \= bi But that’s for European. With American we get inequalities $$AU\_{i+1} \\geq b\_i$$ $$U\_{i+1} \\geq g\_{i+1}$$ The book will mention this g condition, describes the no-arbitrage condition. To solve this system, we have to use what we learned last time about solutions. We can use Jacobi or Gauss-Seidel. Gauss-Seidel is more computationally efficient because you only need one vector. Sometimes Jacobi is faster because Gauss-Seidel goes further, so it might go further in the wrong direction. ### Successive Over Relaxation Recently, 70 years ago, they invented **Successive Over Relaxation (SOR)** which is used a lot in machine learning. We don’t have much ML in the core of financial engineering, so we should be putting some methodology. Bishop had a million ML methods and didn’t explain how. $$U^{(k)} \= U^{(k \- 1)} \+ (U^{(k)} \- U^{(k \- 1)}$$ This illustrates an innovation from an old thing to a new thing. This particular example does nothing. Let’s do this instead. Instead of moving all the way to U^k, let’s move a little bit with ω $$U^{(k)} \= U^{(k-1)} \+ \\omega (U^{(k)} \- U^{(k-1)})$$ If you make it less than 1, than you’re moving a fraction or otherwise you’re moving too much according to Gauss-Seidel. Somehow that doesn’t make sense because we don't have U^(k) yet, so how do we do this? We’ll store the whole Jacobi expression into a single variable y\_j Then you calculate $$U\_j^{(k)} \= U\_j^{(k-1)} \+ \\omega (y\_j \- u\_j^{(k-1)})$$ So basically instead of moving all the way with y\_j, you preserve it a little bit, and you decide how much you move based on ω. You can also change ω at every step, but this is not generally done. This is the SOR method. This is used to solve American options. You don’t need to know the excruciating details, but you should know the big ideas. ### Cubic Splines These were developed to approximate functions Let’s say we observe f: (a, b) \-\> R You don’t know the value of the function, maybe it’s really complicated, you want to approximate it. in regression, you have a bunch of points, and you fit one line. Let’s say our points are all over the place, and a line does not really fit. You could group all the points as some kind of average, between the x and y coordinates, so you get multiple centers of mass for each region. Then you want the curve to go through those points. The initial point distribution is totally irrelevant. It doesn’t have to be a function. What if I have something that goes in circles, or has behavior that doesn’t go in circles. This is the computer science extension which uses B-splines, and conceptually there is no difference. This is the principle. I do an endpoint approximation. We want to approximate f with piecewise polynomials, different polynomials on each segment. We have n \+ 1 knots, these known points. It’s kind of like an anchor point to tie down your curve. In our process, we take $$t\_i \= a \+ \\frac{b-a}{n} i$$ That fractional thing is our Δt. They don’t have to be like this. In general, we can have t\_0, t\_1, … t\_n and it will work the same way. What is the condition? I want to make my curve smooth. I’m going to pick for each interval $$\[t\_i, t\_{i \+ 1}\]$$ we have $$f(t) \= P\_i(t)$$ polynomial What are the conditions on this polynomial? Let’s understand what happens at $$t\_{i-1}, t\_i, t\_{i+1}$$, Well, one condition is that the polynomials have to connect, so $$P\_{i-1}(t\_i) \= P\_i(t\_i) \= f(t\_i)$$ And this becomes two equations for each i. That becomes 2n conditions. Then I have to stitch them. I don’t actually know what the derivative of the polynomial is. If I did know, I could impose the extra conditions that has another 2n conditions, for a total of 4n. But this is only if you know the derivative is, which would make this a lot easier. You need to pick a cubic spline. The realistic condition: $$\\frac{\\partial P\_0}{\\partial t} (t\_i) \= \\frac{\\partial P\_1}{\\partial t} (t\_1)$$ You have 2n \- 2 equations So how do you interpolate? You could take a linear polynomial $$P\_i(t) \= a\_i \+ b\_i t$$ This will clearly not work. You have 2n unknowns and 4n \- 2 equations. The same is true for quadratic. Cubic gives us $$P\_i(t) \= a\_i \+ b\_i \+ c\_i t^2 \+ d\_i t^3$$ Now we have 4n unknowns and 4n \- 2 equations. So how do we fix this? There are n \+ 1 points. There’s a very clear order unless you somehow define it. And that determines how you stitch the functions. There’s an infinite number of solutions by fixing the two missing equations. In traditinoal splines that come from statistics, the way it works is that there’s two possibilities. There can be a free boundary problem where the $$S\_0’’ t(0) \= S\_nn’’ t(n) \= 0$$ When you say the second derivative is 0, that’s the maximum of the first derivative. It starts with the largest slope possible, Then the second one is the clamped boundary $$S\_0’ (t\_0) \= f(t\_0)$$ $$S\_n’ (t\_n) \= f’(t\_n)$$ Because you don’t know the derivative, you constrain the curve to have a certain fixed slope. You constrain the first one to start at a certain angle. Then how do you solve this? It’s in terms of the functions and their derivatives. I plug in these points and gets an equations in a\_i, b\_i, c\_i, and d\_i, and then the next parameters. Basically I have four end parameters, and since these equations are relative to each other, it becomes a sparse system of equations. It’s not tridiagonal, but it does have a system. Then you solve it with Jacobi, Gauss-Seidel, SOR, etc. when you solve it in R. These are very useful for smoothing approximations of curve. I have used them in my lecture earlier in the pricing the implied vol surface paper. This whole paper is inspired by two dimensional cubic splines. Basically, the Derman people created a bullshit local volatility problem. I thought about it and I realized, if you give me the option prices, and get the IV, and then fit a curve through these things. R had just created the 2D approximation and just used this. It’s basically the same as just described, but you do it with a line instead of points, but I don’t remember how and it’s pretty complicated. There are many many applications to all of these numerical methods. I learned about this realized volatility surface from these people, reading about this other method of approximation, and putting them together. That’s why it’s very important to learn approximation methods. There’s many people that start the task, cut government waste, and then you go in there and have no clue what you’re doing. Method, knowing what to do is the most important thing. Either you get videos from FOX that are BINGQILIN ## FX Volatility Smiles and Uncertainty - URL: https://sharifhsn.dev/blog/fx-volatility-smiles-and-uncertainty/ - Structured data: https://sharifhsn.dev/api/posts/fx-volatility-smiles-and-uncertainty.json - Description: The next FE-635 notes explain why a single Black–Scholes volatility is not enough for an FX market. Discrete hedging creates P&L even when the model's inputs are correct, and volat… - Date: 2025-03-10 - Exact published timestamp: 2025-03-10 - Topics: FX, Volatility, Smiles - Categories: FX - Source: FE-635 \| Risk Engineering - Source URL: None The next FE-635 notes explain why a single Black–Scholes volatility is not enough for an FX market. Discrete hedging creates P&L even when the model's inputs are correct, and volatility is uncertain rather than fixed. Options with different strikes imply different volatilities. The resulting smile or skew is a market summary of tail demand, quotation conventions, and the limits of the log-normal model. Calibration therefore fits a surface of prices or implied volatilities instead of forcing every option through one number. The practical implication is clear: risk must be revalued under shocks to both spot and the volatility surface. A delta-only report can miss the largest loss when the smile moves with the underlying. ## Hull–White Interest Rate Trees - URL: https://sharifhsn.dev/blog/advanced-derivatives-week-06/ - Structured data: https://sharifhsn.dev/api/posts/advanced-derivatives-week-06.json - Description: This will be focused on numerical methods, next week we will have more examples. - Date: 2025-03-06 - Exact published timestamp: 2025-03-06 - Topics: Fixed Income, Hull-White Model, Trinomial Trees, Interest Rate Models - Categories: Fixed Income - Source: Advanced Derivatives - Source URL: None ## Numerical construction of the trinomial tree for the Hull-White model This will be focused on numerical methods, next week we will have more examples. What is the difference between interest rate risk and stock price risk? One of the most important differences is that the discount rate will vary from node to node. We will look at an example: We’ll assume that the payoff of a derivative is 100(r-0.11)\_+ We also know the probability of going up, staying and going down as 0.25, 0.5, 0.25 ![Hand-drawn Hull-White trinomial tree from the source notes.](/static/img/Advanced Derivatives-week-06-hull-white-tree.png) Hull-White uses an alternative branching process. There is the typical (up med down) then (UP up med), (med down DOWN). $$ dr = [\\theta(t) - ar] dt + \\sigma dz $$ So what is the procedure? First we have to assume \\(\\theta(t) = 0\\) And the start is \\(r(0) = 0\\) The construction of this trinomial tree should match the expectation and variance of the process? Basically we run the simulation based on the parameters. Then we determine \\(\\theta(t)\\) so it matches the initial term structure. Let’s look at the stages: ![Handwritten tree-construction steps from the Week 6 source notes.](/static/img/Advanced Derivatives-week-06-tree-construction.png) Then we set \\(\\Delta R = \\sigma \\sqrt{3 \\Delta t}\\) This particular equality is chosen for numerical convergence. We need to determine which branching method applies at each node: (regular, up, down). Let each node be \\((i, j)\\) where \\(t = i \\cdot \\Delta t\\) and \\(R* = j \\Delta R\\) So this is the position in the state space. As for indices, i is related to time, it is positive. j can be positive or negative because Hull-White allows negative interest rates. switching the branching scheme is dependent on the index j **when a \> 0** switch from branching *regular* to *down* for sufficiently large positive j. switch from branching *regular* to *up* for sufficiently large negative j. Let \\(j_{\\max}\\) and \\(j_{\\min}\\)for the value where you switch Hull and White (posted on Canvas) show that the probabilities are always positive *if* \\(j_{\\max}\\) is equal to the smallest integer \> \\(\\frac{0.184}{a \\Delta t}\\) $$ j_{\\min} = -j_{\\max} $$ The point where we change branching is both dependent on the speed of mean reversion and the length of the time step. Therefore we can say that the probabilities \\(p_u\\), \\(p_m\\), and \\(p_d\\) are chosen to match expected change and variance of change in R\*. We will solve for these unknowns with equations. What are those equations? That depends on the branching strategy: *regular*: $$ p_u \\Delta R - p_d \\Delta R = -aj\\Delta R \\Delta t $$ $$ p_u \\Delta R^2 + p_d \\Delta R^2 = \\sigma^2 \\Delta t + a^2 j^2 $$ $$ p_u + p_m + p_d > 1 $$ So this is your system of equations. And by setting \\(\\Delta R = \\sigma \\sqrt{3 \\Delta t}\\), our system becomes $$ p_u = \\frac{1}{6} + \\frac{1}{2}(a^2 j^2 \\Delta t^2 - aj \\Delta t) $$ $$ p_m = \\frac{2}{3} - a^2 j^2 \\Delta t^2 $$ $$ p_d = \\frac{1}{6} + \\frac{1}{2} (a^2 j^2 \\Delta t^2 + aj\\Delta t) $$ Note that these are all dependent on index j. We will look at an example of this. Similarly, for *up* $$ p_u = \\frac{1}{6} + \\frac{1}{2}(a^2 j^2 \\Delta t^2 + aj \\Delta t) $$ … I didn’t write them down in time. And for *down* $$ p_u = \\frac{7}{6} + \\frac{1}{2}(a^2 j^2 \\Delta t^2 - 3aj\\Delta t) $$ $$ p_m = -\\frac{1}{3} - a^2 j^2 \\Delta t^2 + 2aj\\Delta t $$ $$ p_d = \\frac{1}{6} + \\frac{1}{2}(a^2 j^2 \\Delta t^2 - aj \\Delta t) $$ Let’s take values of a zero curve 0.5 1 1.5 2 2.5 3 3.43 3.824 4.183 4.512 4.812 5.086 Building the first tree, we set our **vertical spacing** to the special numerical value \\(\\sigma \\sqrt{3\\Delta t}\\) Then what is the value of j in our example? $$ \\sigma = 0.01 $$ Then after we construct the tree, we have to displace the nodes to match the original term structure. **STEP 2** Convert this R\* tree into a tree for R, where we displace the nodes so that we match the initial term structure. Our displacement function is $$ \\alpha(t) = R(t) - R*(t) $$ These can be calculated from the \\(\\theta(t)\\) expression. Recall that $$ \\theta(t) = F_t(0, t) + a F(0, t) + \\frac{\\sigma^2}{2a}(1 - e^{-2at}) $$ Note $$ dR = [\\theta(t) - aR] dt + \\sigma dW $$ $$ dR* = -aR* dt + \\sigma dW $$ So therefore if we substitute in, $$ d\\alpha = [\\theta(t) - a\\alpha(t)] dt $$ Then we can rearrange and integrate to get $$ \\alpha(t) = F(0, t) + \\frac{\\sigma^2}{2a^2} (1-e^{-at})^2 $$ for infinitesimally small values. So this is kind of a continuous function. In this case, we need finite Δt to match the term structure exactly. So it’s a little more simple We define \\(\\alpha_i\\), which corresponds to each time step on the tree, $$ \\alpha(i \\Delta t) = R(i \\Delta t) - R*(i \\Delta t) $$ Here we’re interested in the discrete points on the tree. Then the overall procedure here is that we are given some rates on the zero curve, but you can have multiple rates on the time step. So you have to price, and look at the value given by the market, and the value given by the tree. Then we can define the present value of the security \\(Q_{i, j}\\) that pays $1 if node (i, j) is reached and $0 otherwise, so we can add up the probabilities and discount them. And we can calibrate α and Q to match the initial term structure. So how do we start? At 0, by definition, we pay $1 $$ Q_{0, 0} = 1 $$ And the displacement of the initial node is chosen to give the right price for a ZCB maturing at Δt. Since our zero rate is 3.824 for a one-year rate, we can set α\_0 to be that. In an exercise, if you were to take Δ to be 0.5 it would be that rate. And then \\(\\alpha_1\\) should be chosen in such a way to be a ZCB that matures in two years. Then next we calculate \\(Q_{1,1}, Q_{1,0}, Q_{1,-1}\\). i.e., at time i=1, there are three possibilities for j. We use the probabilities, for \\(Q_{1,1}\\) it’s 0.1667. And then the discount factor is the value of R, in this case 3.824%. we get the three values 0.1604, 0.6417, and 0.1604. And then we can use it to calculate \\(\\alpha_1\\) which is chosen to match the right price of a ZCB maturing at 2Δt. We have the rate from the market, so we can price the bond with that rate. Such bond at 2Δt, you will discount with the rate applicable per branch, which is 1.732%. What we will see in this displaced tree, will be this original value 1.732% discounted. $$ e^{-(\\alpha_1 + 0.01732) \\times 1} $$ Then for node C, the rate 0 will be ignored so we just use α\_1 Then having the price of the bond at nodes B, C, and D, we can calculate the price of the node. $$ A = Q_{1,1} e^{-(\\alpha_1 + 0.01732)} + Q_{1,0} e^{-\\alpha_1} + \\ldots $$ From the initial term structure, we can get the ZCB for two years as the “market price” \= 0.9137. And then we substitute in the Q values to get the total formula for the adjustment \\(\\alpha_1\\) $$ \\ln\\left[\\frac{Q_{1,1}e^{R_{1,1}} + Q_{1,0} + \\ldots}{\\text{P(0, 2)}}\\right] $$ **Here is the formal approach**. If we have \\(Q_{i,j}\\) already determined for \\(i \\leq m\\) where m is the final step of the tree. Then BIG LONG FORMULA on page 18\. If you have the transition probability matrix, it’s an iterative approach where you determine α, then Q, then α, then Q. ## Recombining binomial tree ## Crank–Nicolson and the Heston Model - URL: https://sharifhsn.dev/blog/computational-methods-week-06/ - Structured data: https://sharifhsn.dev/api/posts/computational-methods-week-06.json - Description: Don’t use AI, you won’t learn anything. - Date: 2025-03-04 - Exact published timestamp: 2025-03-04 - Topics: Computational Methods, Crank-Nicolson, Linear Systems, Heston Model, Finite Differences - Categories: Computational Methods - Source: Computational Methods in Quantitative Finance - Source URL: None ## AI Rant Don’t use AI, you won’t learn anything. And it’s not really intelligent, it just takes things from the Internet. ## Crank-Nicholson Finite Difference Methods 7.4 in the book. There are two different general finite difference schemes, which are about which points you use in the grid. In the explicit scheme, you go from many points to one point, so you can find out each point from the points in the boundary, kind of in a trinomial tree way. In the implicit scheme, you go from one point to three points, which allows you to solve the systems, and find all the points altogether. Crank-Nicholson is a variation of the implicit finite difference that is better than both, because it is more complicated. We have the points. Remember that \\(i\\) is \\(\\Delta t\\), and \\(j\\) is \\(\\Delta x\\). We need three points for each derivative. We will actually use the average of the points to calculate the x-type derivative. This is the definition of ad derivative in a numerical sense: $$ \\frac{\\partial u}{\\partial t} = \\frac{u(i + 1, j) - u(i, j)}{\\Delta t} $$ For this one, we use the top yellow minus the bottom yellow. It’s a bit tricky because it’ saveraged, the point is to use both at the same time. Then you divide by the interval between them. $$ \\frac{\\partial u}{\\partial x} = \\frac{\\tfrac{1}{2}(u(i, j+1) + u(i + 1, j + 1)) - \\tfrac{1}{2}(u(i, j - 1) + u(i +1, j - 1))}{2\\Delta x} $$ Then the second derivative is horrendous. It works by using the two second derivatives and reducing to the value of the topmost plus the value of the bottom most minus twice the middle $$ \\frac{\\partial^2 u}{\\partial x^2} = \\frac{\\tfrac{1}{2}(u(i, j+1) + u(i + 1, j+1)) - (u(i, j) + u(i + 1, j)) + \\tfrac{1}{2}(u(i, j-1) + u(i + 1, j - 1))}{\\Delta x^2} $$ We have all these derivatives in the equation, but we will also use the u. We will use u\_i for the last term, which is just u(t, x) for easier usage. I’m usinthe mid points, the yellow points, all over the place. If I plug them into the equation, what do I get? $$ \\frac{\\partial u}{\\partial t} + \\nu \\frac{\\partial u}{\\partial x} + \\frac{1}{2}\\sigma^2 \\frac{\\partial^2 u}{\\partial x^2} - ru = 0 $$ where \\(\\nu = r - \\tfrac{\\sigma^2}{2}\\) Then, plugging in our finite differences, it’s $$ \\frac{u(i + 1, j) - u(i, j)}{\\Delta t} + \\nu \\frac{u(i, j +1) + u(i + 1, j + 1) - u(i, j - 1) - u(i + 1, j - 1)}{4\\Delta x} + \\frac{1}{2} \\sigma^2 \\frac{u(i, j + 1) + u(i + 1, j + 1) - 2(u(i, j) + u(i + 1, j)) + u(i, j - 1) + u(i + 1, j - 1)}{2 \\Delta x^2} - r\\frac{u_{ij} - u_{i+1, j}}{2} = 0 $$ If this doesn’t work, please follow up because it might not work I will isolate all terms that have terms u(i) because those are unknown, I already know i \+ 1\. So I should factor out these three terms here to get our coefficients $$ u(i, j + 1)(\\frac{\\nu}{4\\Delta x} + \\frac{\\sigma^2}{4\\Delta x^2}) - u(i, j)(\\frac{1}{\\Delta t} + \\frac{\\sigma^2}{2\\Delta x^2} + \\frac{r}{2}) + u(i, j - 1)(-\\frac{\\nu}{4\\Delta x} + \\frac{\\sigma^2}{4\\Delta x^2}) $$ The point of being careful about this is that this is what gives us our A, B, and C. So it becomes Ax \+ By \+ Cz \=... And now this is the nightmare you have to move to the other side, and make sure to change the sign. $$ = -\\frac{1}{\\Delta t} u(i + 1, j) - \\frac{\\nu}{4\\Delta x}(u(i + 1, j + 1) - u(i + 1, j - 1) - \\frac{\\sigma^2}{4\\Delta x^2} (u(i + 1, j + 1) - 2u(i + 1, j) + u(i + 1, j - 1)) - \\frac{r}{2} u_{i+1, j} $$ Now I have everything I need to do the pseudocode. ### Pseudocode Initialize \\(\\Delta x\\) values. Then I get the payoffs on the right most boundary based on \\((e^{x+ n\\Delta x} - K)_+\\) or whatever. That is my values for u(n, N), u(n, N \- 1)...u(n, \-N) These equations supposedly work from top to bottom. Because I can’t go to infinity. I need to add two boundary numbers. And all of these numbers should recursively go backwards. Then you will end up with the tridiagonal system with A B C, 0 A B C, 0 0 A B C, etc. Then you will also need the \\(\\lambda\_u\\) and \\(\\lambda \-L\\) to finish the vector. The system has the very same form as the implicit finite difference method. The only difference is that the last b vector is solved in a complex way. ### Stability How is this method stable with floating point errors? You are dividing by A \+ B, which as long as it’s not tiny, it won’t explode. ## Methods to Solve Linear Systems of Equations Imagine you have to solve Ax \= b. Jacobi says, let’s do a trick. (This only works for nonzero diagonal systems). $$ A = \\begin{bmatrix} a_{11} & a_{12} \\\\ a_{21} & a_{22} \\end{bmatrix} $$ The diagonal cannot be zero. You will take the lower triangle of the matrix L, the diagonal D, and the upper triangle U. This allows me to do $$ (L + D + U)x = b $$ You would normally take the inverse to get $$ x = A^{-1} b $$ But Jacobi was a finance guy. When you try to invert 2x2 it’s okay, 3x3 is painful, and 4x4 is unbelievable. So before they had calculators, he wanted to invert something easier. $$ Dx = b - (L + U)x $$ Basically only inverting the part I know how to invert easily, since inverse diagonal is just the inverse of all its elements. $$ x = \\frac{1}{D}(b - (L + U)x) $$ That might seem useless, but then we can do this: Let’s start with some random x^0, maybe all zeroes, then plug this into the equation, and take the following recurrence: $$ X^(k) = D^{-1} (b - (L + U) X^{(k - 1)}) $$ This looks ugly in matrix form, but it’s more easily expressed as: $$ a_{11} x_1 = b_1 - a_{12} x_2 - \\ldots - a_{1n} x_n $$ $$ x_2 = \\frac{1}{a_{22}}(b_2 - a_{21} x_1 - a_{23} x_3 - \\ldots - a_{2n} x_n) $$ So you just move everything to the other side. We’ll take \\(x^{(0)}\\) to be all zeroes. Then you take the result in the next x, and recur it, and Jacobi proved that it will converge, AS LONG AS the diagonal elements are nonzero because notice that we have to divide by them. There’s another generalization of this that will be useful when we talk about American options pricing. ## Gauss-Seidel This is the PPO of the Jacobi, where it’s a very small change from the Jacobi. $$ x_1^k = \\frac{1}{a_{11}} x_2^{(k - 1)} + a_{13}x_3^{(k-1)} \\ldots + a_{1n} x_n^{(k-1)}) $$ But if you notice in x\_2, we already calculated the previous recurrence. And this does two steps in 1, and it has faster convergence. This is a very common idea called the fixed point method. Whenever you solve equations, you look for things that are fixed. If you calculate a derivative, the reason is because it is the maximum. However, when you get close to the maximum, and you get bigger, then you get stuck. The gradient method goes up and up until you can’t go anymore. You need to see this pattern everywhere, because it’s very important. The human brain is stupid. ## Finite Difference for Heston This is pretty tough, so we won’t get all the details. There are some convergence issues with the explicit finite difference method. You need the conditions that \\(\\Delta x \\geq \\sigma \\sqrt{3\\Delta t}\\) and \\(N = n\\). Also, the order of convergence is \\(O(\\Delta x^2 + \\Delta t)\\). So you want \\(\\Delta x\\) to be as small as possible, to make convergence faster, so effectively the optimal is that it equals this. We will pick an n, and then the following values fall out of it: $$ \\Delta t = T/n, \\Delta x = \\sigma \\sqrt{3\\Delta t}, N = n $$ There’s nothing to choose. But you have some more freedom for IFD (Implicit Finite Difference) and Crank Nicholson. We have the same order of convergence $$ O(\\Delta x^2 + \\Delta t) $$ but this is unconditionally convergent, so it doesn’t matter what values you pick, it doesn’t have this condition about being greater than whatever. Let’s say that you want to be within epsilon of the true number. So how do I make sure that I’m within? Technically you can’t, because big O has constants associated with it, but we might take $$ \\Delta x^2 + \\Delta t = \\epsilon $$ Because there is no relationship between them, you can pick anything for either. Let’s say \\(\\Delta x^2 = \\tfrac{\\epsilon}{2}\\) and the same for \\(\\Delta t = \\tfrac{\\epsilon}{2}\\). This gives you a known \\(\\Delta t\\) which will give you n. But \\(\\Delta x\\) is not fixed. But how far should I go? (N). Because this process is Black-Scholes, we know that the spread at time T is lognormal, so \\(X_T \\sim N(x_0, \\sigma^2 T)\\) So we need N such that \\(N \\Delta X > 3.5 sigma \\sqrt{T}\\) We know that most values are within 3 standard deviations, I do 3.5 just in case, but generally it doesn’t matter. For the Crank-Nicholson, we have $$ O(\\Delta x^2 + \\frac{\\Delta t^2}{2}) $$ So we handle it in kind of a similar way to IFD. How you split it is your choice. Last thing is about Calculating Greeks: These finite difference methods are useful for calculating these derivatives. You are approximating derivatives, and the Greeks are derivatives. Gamma, theta, vega, are instilled in the u-values. Delta x is not the stock price, so you actually need to take e^x\_0 For Delta: $$ \\frac{u_{0, 1} - u_{0, -1}}{e^{x_0 + \\Delta x} - e^{x_0 - \\Delta x}} $$ And similar for Gamma, we will need all three points, in the same way as the other second derivative was described. The solution to the PDE is the same as the price of the option here, because that’s how we express the option. So we use the finite difference to numerically solve the PDE. The best way to solve a PDE is analytically, then you’re done. But, I can count on my hand how many PDEs you can actually solve, just three. So in general, if you write anything (and the heat equation is one where there is a solution) it’s complex and you have to approximate it. There is no formula for the Heston model. It involves a characteristic and inverse characteristic function, which is the probabilistic version of a Fourier transform. And there are not that many known analytical solutions for the Fourier Transform. So you do the discrete Fourier transform and the inverse discrete Fourier Transform. That’s how the solution for the Heston model is expressed. The ultimate way to solve for the Fourier is to do this finite difference method for the Heston model. In order to solve the Heston model, we need to use a bunch of parameters. r is the reglar risk free interest rate. $$ dS_t = (r - q)S_t dt + \\sqrt{y_t} S_t dW_t $$ This is the same as Blck-Scholes but \\(\\sigma\\) is replaced by this other stochastic process \\(y_t\\). And the parameters are $$ dY_t = \\kappa (\\bar{y} - y_t) dt + \\sigma \\sqrt{y_t} dW^2_t $$ where the two Brownian motions are correlated with \\(\\rho\\) This is a Cox Ingersoll Ross process (or Ornstein Uhlbeck) Let’s consider Black-Scholes: one reason it’s nice is because you can do this logarithm transformation like this: $$ dS_t = rS_t dt + \\sigma S_T dW_t $$ $$ dR_t = (r - \\tfrac{\\sigma^2}{2}) dt + \\sigma dW_t $$ So you might use this for interest rates, but actually it’s bad. So you usually use this Vasicek (or Ornstein Uhlerbeck) model: $$ dR_t = \\kappa(r - R_t) dt + \\sigma dW_t $$ where r is the long term interest rate what the Fed does an model the short term interest rate. The problem is that it can be negative. So this was created in 1978 and it was rejected immediately. Then in 2008, interest rates went negative and now Vasicek is what people actually use. And this can actually be solved. If you’re taking a qualifying exam you should know how to do this. Then you have the Cox-Ingersoll-Ross model where you take the mean reverting feature and then add a square root term $$ dR_t = \\kappa (r - R_t) dt + \\sigma \\sqrt{R_t} dW_t $$ That R\_t makes the variance dependent ont he value of the proess. So when it goes close to 0, the volatility goes to 0 and the mean reverting term moves the process more to the mean. And this is the process used everywhere. And you can see that this is the same dynamic for the Heston model. However, unlike the CIR, this cannot be solved, does not have analytical solution. One way to solve this is with transformations, which we will do after the midterm. Another way is to come up with the PDE and solve the PDE with our other solutions. Writing the PDE is not easy, you have to make a lot of no-arbitrage arguments. ### Heston Option Price v is actually the variance, it is the value for the Y process. $$ u(t, S, v) $$ You can derive the equation, but I will write it down like I’m God. q is the dividend thing. It’s a general thing for if you have a continuously paid dividend. \\(\\theta\\) is the \\(\\bar{y}\\) in the heston model. $$ \\frac{\\partial u}{\\partial t} = \\frac{1}{2} vS^2 \\frac{\\partial^2 u}{\\partial S^2} + \\rho \\sigma v S \\frac{\\partial^2 u}{\\partial S \\partial v} + \\frac{1}{2} \\sigma^2 v \\frac{\\partial^2 u}{\\partial v^2} - ru + (r - q) S \\frac{\\partial u}{\\partial S} + \\kappa (\\theta - v) \\frac{\\partial u}{\\partial v} $$ In order to get this nice equation, we also take a change of variable where \\(t = T - \\tau\\) so this goes from a terminal problem to an initial problem. So I actually know the solution $$ u(0, S, v) = (S - K)_+ $$ The reason why it’s like this is because in the solvers you have, they are designed to work as an initial value problem. For all these terms, you need to substitute them and do the derivative stuff. Everything else we’ve seen, but the mixed derivative is a little weird. It’s in the book, though. The bigger deal is that these coefficients depend on the location you are. They depend on where you are, they depend on S and v. So in the equations, every time the coefficients are dependent on the grid. When you move back (or forward in this case) you have to keep track of your grid points and change your coefficients based on where you are. There are two approaches: **Fixed Grid**: You will have three dimensional grid where you have to keep track of stock, time, AND variance. This is on page 190-191 of the book. You read the corresponding points, you plug them into the equation, and the solution is solved in the same way. The scheme here is an explicit scheme. The implicit scheme is absolutely horrendous for this. **Smart Grid**: Instead of using a rectangular grid, you will put your points in non-equally spaced space. The coefficients will become constant if you do it in this way. The coefficients are easy to solve. ## Remember This There are two different types of finite difference, EXPLICIT AND IMPLICIT and implicit has some variations. It’s about discretizing the derivatives. You have to look at every equation you discretize, which points you bring into the equations, because the equation is a relationship between these points. It’s all about the points you pick. If your boundary is not a straight line, say a curve, then you want to form a relationship that has three points you know, a point you don’t know. That’s the explicit scheme. Which is easier, but it has some drawbacks and some limitations, as descried relating to convergence. Implicit gives you a system of equations, which will be linear, and eventually you will solve it. ## Research I wrote 3 or 4 papers on this. All of the equations, if you look, it’s like $$ \\frac{\\partial u}{\\partial t} = L u $$ where the operator L contains all this second derivatives, mixed derivatives, whatever. We created a way to solve this, in a way that is parallel to the Jacobi idea. You start with some u0 and plug it in here, take the derivative of the function, and then you get the derivative, which you can replug and recur. Their methodology does both: it moves through the tree and recurs at the same time. It’s slow, but it works. ## Finite-Difference Sketches - URL: https://sharifhsn.dev/blog/computational-methods-finite-difference-sketches/ - Structured data: https://sharifhsn.dev/api/posts/computational-methods-finite-difference-sketches.json - Description: A handwritten finite-difference sketch with A/B/C labels; surrounding equation notation is unclear in the source render. - Date: 2025-03-04 - Exact published timestamp: 2025-03-04 - Topics: Computational Methods, Finite Differences, Grid Sketch, Numerical Methods - Categories: Computational Methods - Source: Handwritten FiniteDifference.xopp and matching autosave, byte-for-byte identical by SHA-256; Drive-modified 2025-03-04. Approximate date within the Computational Methods in Quantitative Finance Docs history cluster (2025-01-21 through 2025-05-15). - Source URL: None The week 5 and 6 notes contain the broader finite-difference grid and explicit/implicit, Crank–Nicolson, and Heston derivations. This page is kept as a brief sketch supplement rather than repeating those derivations. The hand-drawn diagram has the readable labels `A`, `B`, and `C` beside several tall, curved strokes. Other point labels and the nearby coefficient and equation notation are **unclear in the source**. ## FX Discrete Hedging and Greeks - URL: https://sharifhsn.dev/blog/fx-discrete-hedging-and-greeks/ - Structured data: https://sharifhsn.dev/api/posts/fx-discrete-hedging-and-greeks.json - Description: The FE-635 notes move from an ideal continuous hedge to the discrete hedges used in practice. A delta hedge removes the first-order response to an FX move at one instant; between r… - Date: 2025-03-03 - Exact published timestamp: 2025-03-03 - Topics: FX, Greeks, Discrete Hedging - Categories: FX - Source: FE-635 \| Risk Engineering - Source URL: None The FE-635 notes move from an ideal continuous hedge to the discrete hedges used in practice. A delta hedge removes the first-order response to an FX move at one instant; between rebalances, spot and volatility can move and leave a residual. The workbook exposes the main sensitivities directly: delta, gamma, vega, and domestic and foreign rho. Gamma measures how quickly delta changes, vega measures sensitivity to volatility, and the two rho functions separate the currencies' discounting effects. The lesson is operational as much as mathematical. A hedge report should state the rebalance frequency, quote direction, and whether the risk is measured per unit of foreign currency or in domestic dollars. ## Constant Maturity Swaps and Short-Rate Models - URL: https://sharifhsn.dev/blog/advanced-derivatives-week-05/ - Structured data: https://sharifhsn.dev/api/posts/advanced-derivatives-week-05.json - Description: Timing and quanto adjustments, we also looked at nonstandard swaps, LIBOR for LES? Another kind of nonstandard swap is… - Date: 2025-02-27 - Exact published timestamp: 2025-02-27 - Topics: Fixed Income, Constant Maturity Swaps, Short Rate Models, Vasicek, Cox-Ingersoll-Ross, Hull-White - Categories: Fixed Income - Source: Advanced Derivatives - Source URL: None ## Continuing Adjustments Timing and quanto adjustments, we also looked at nonstandard swaps, LIBOR for LES? Another kind of nonstandard swap is… ## Constant Maturity Swaps (CMS) The definition of a CMS is **an interest rate swap where the floating rate equals the swap rate for another reference swap for a certain life** If you want to write an example, here you can say floating payment on a CMS can be made every six months at rate equal to 5-year swap rate. The frequency of this CMS rate must be smaller than the reference rate. Usually the swap rate is equal to a previously observed swap rate. Furthermore, if we assume that we have reset dates \\(t_0, t_1, t_2 \\ldots\\) And payment dates: \\(t_1, t_2, t_3 \\ldots\\) We will also consider L \= notional principal What is going to be the floating payment? At \\(t_{i+1}\\), this would be the swap rate \\(s_i\\) applied to the principal for a period \\(\\tau_i\\) $$ \\tau_i L s_i $$ where \\(\\tau_i = t_{i+1} - t_i\\) Now we have to do some adjustments. The convexity adjustment is, the frequency of these coupons is going to be the same as the bonds. We will also have to do a timing adjustment, a combination of these two. If you apply these to another currency, then we will also have a quanto adjustment. Let’s say \\(y_i\\) is the forward value of the swap rate. By using a convexity/timing adjustment. We can say the **realized swap rate** is assumed: $$ y_i - \\frac{1}{2} y_i^2 \\sigma^2_{y_i} t_i \\frac{G_i''(y)}{G_i' (y)} - \\frac{y_i \\tau_i F_i\\rho_i \\sigma_{y,i} \\sigma_{F,i} t_i}{1 + \\bar{T}_i \\tau_i} $$ where \\(\\sigma_{y, i}\\) is the volatility of the forward swap rate \\(F_i\\) is the current forward interest rate between \\(t_i\\) and \\(t_{i+1}\\) \\(\\sigma_{F, i}\\) is the volatility of the forward rate \\(G_i(x)\\) this is basically the price of a bond which pays at time i, the same amount as the reference swap. The forward rate will be implied by the “Kaplan?” prices. \\(\\rho_i\\) the correlation can be estimated from historical data. ### Example 6 year CMS swap (life of swap is 6 years) 5 year swap rate is received, fixed rate of 5% is paid on notional $100M exchange of payments is semiannual, on both 5-year underlying swap and on CMS swap. The exchange rate on the payment date is determined from the swap rate on the previous payment date. The term structure is flat at 5% per annum with semiannual compounding. (that’s where the timing adjustment comes in) The volatility of the options on five-year swaps have 15% IV (\\(\\sigma_{y, i}\\)) and in this simplified problem we’ll say caplets with 6 month tenor have a 20% IV. The correlation \\(\\rho_i\\) between each cap rate and each swap rate is 0.7 The fixed rate \\(y_i\\) is 0.05. This is semiannual payment, so \\(\\tau_i = 0.5\\). Given volatility \\(\\sigma_{F, i} = 0.20\\). Then the forward rate is also \\(F_i = 0.05\\) from the term structure, and the correlation given is \\(\\rho_i = 0.7\\). Then what is the relevant function G in this case? We have the reference rate which is a 5 year swap rate, which has a semiannual payment. We can express this as a bond, so a sum of coupons. The 2.5M is the coupon payment (halved $$ G_i(x) = \\sum_{i=1}^{10} \\frac{2.5}{1 + \\tfrac{x}{2})^i} + \\frac{100}{(1 + \\tfrac{x}{2})^10} $$ Then the derivative is $$ G_i'(y_i) = -437.603, G_i'' (y_i) = 2261.23 $$ Then the total convexity/timing adjustment is, \\(0.0001197t_i\\) We’re going to have different values, so for example the 5 year swap rate t \= 4 should be assumed to be 5.0479% instead of 5%. Therefore the net cash flow at t \= 4.5 should be 0.5 \* 0.000479 \* 100M. Then we can calculate for all these cash payments, the value of the CMS ## Models of the Short Rate There are different specifications of such models. The models presented so far in our previous discussions make the assumption that the probability distribution of variables (interest rates, bond prices, …) are log normally distributed. That allows us to use Black’s model. There are some issues with this kind of approach. - These models do not provide a description on how the interest rates evolve in time, just the underlying distribution. - Cannot be used to value American style options. In this lecture, we will discuss the **“term structure model”**. Such interest rate model will describe the evolution of all zero-coupon interest rates, a little more general. We will focus on some classical models. The behavior of the **“short-rate”** or instantaneous short rate. This rate is considered in a risk-neutral world. Basically, that means in a very short period of time, between t and \\(t + \\Delta t\\), investors earn \\(r(t) \\Delta t\\), the risk-free interest rate. For example, the dollar money market account is a security that is worth $1 at time 0 and earns the instantaneous risk-free rate r at any given time. This r may be stochastic. Furthermore, if g equals the money market account, then the process \\(dg = rg dt\\). Here we can see that the r the risk-free rate can be stochastic, so this will be the drift of g being stochastic, with volatility being 0\. We saw last lecture that \\(\\frac{f}{g}\\) is a martingale in a world where the market price of risk is 0. Therefore the expected value is $$ f_0 = g_0 \\hat{\\mathbb{E}}[\\frac{f_T}{g_T}] $$ Then if you take g as the money market account, then by definition $$ g_0 = 1 $$ And for \\(g_T\\)? You might expect the e^rt general, but because this is stochastic, we express this as $$ g_T = e^{\\int_0^T r dt} $$ Therefore, the expectation of f at T must be discounted $$ f_0 = \\mathbb{E}[f_T e^{-\\int_0^T r dt}] $$ You could also express this as $$ f_0 = \\hat{E}[e^{-\\bar{r}T} f_T] $$ where \\(\\bar{r}\\) is the average value of r. ## Interest Rate Derivative Valuation A regular way to value an interest rate derivative, typically these will follow these steps: 1. Simulate short-term interest rate paths. (in some cases we will have closed-form solutions, but generally we simulate) 2. Calculate expected payoff on these paths. 3. Discount at average value of short rate on the sampled path. In general, the value at time t of an interest rate derivative is based on providing a payoff \\(f_T\\) at time T $$ \\hat{\\mathbb{E}}[f_T e^{-\\bar{r}(T - t)}] $$ Basically you calculate the value of such derivatives between t and T. Furthermore, if you refer back to \\(P(t, T)\\) for the price function at time t of a ZCB with payoff of $1 at time T, we can say that $$ P(t, T) = \\hat{\\mathbb{E}} [e^{-\\bar{r}(T - t)}] $$ Now we will make connection to the interest rate. Generally, if R(t, T) is continuously compounded interest rate at time t, for a term (T \- t), then we can say $$ P(t, T) = e^{-R(t, T)(T - t)} $$ then we can rearrange as $$ R(t, T) = -\\frac{1}{T - t} \\ln P(t, T) $$ Then finally we will arrive at the relationship $$ R(t, T) = -\\frac{1}{T-t} \\ln \\hat{\\mathbb{E}}[e^{-\\bar{r}(T-t)}] $$ This equation is what enables the calculation of the term structure of the interest rate at any given time. Based on the description of the risk-free rate, which may be stochastic, can help you describe the entire curve. All the models we discuss will be models of this short rate, there are some models of the instantaneous forward, but generally they will be instantaneously spot. These classical models give descriptions of stochastic processes that the short rate will have this particular expression. And this is the zero curve ## Equilibrium Models Rendleman & Bartter, Vasicek, Cox Ingersoll & Ross These models start with some assumptions about the economic variables, then the variable process from the short rate r. This will explore what the process for r implies about the bond prices and option prices. This model might be one factor, or two factor. We will look now just at one factor models. In a one factor equilibrium model, the process for r involves only one source of uncertainty. $$ dr = m(r) dt + s(r) dW $$ This general specification, m(r) is the instantaneous drift, and s(r) is the instantaneous volatility (stdev). For these models, m(r) and s(r) are *assumed* to be a function of r, but *independent* of time. This is different from the no-arbitrage models. Some of these models, historically, have been proposed. Some of them are not necessarily very good (lol). ### Rendleman and Bartter $$ dr = \\mu r dt + \\sigma r dW $$ Where \\(\\mu\\) and \\(\\sigma\\) are constant, it basically means that r follows GBM. But this is bad because r doesn’t follow GBM. These interest rates have some mean-reverting behavior, so that has to be included somehow. When r is high, it will tend to have negative drift, and when it’s low, it will tend to have positive drift. Economic argument, high r will cause the economy to slow down, less borrowing, and vice versa. These models can also be calibrated, will be discussed next time. ### Vasicek $$ dr = a(b - r) td + \\sigma dW $$ where a, b, σ are constants. Here we can see the mean reverting component. In the long-run, the value, when r is small, it will be a positive number and increase, and vice versa. a is the speed of mean reversion, b is the mean. This model has an analytical solution. Vasicek shows that if $$ P(t, T) = \\hat{\\mathbb{E}}\\left[e^{-\\bar{r}(T-t)}\\right] $$ then the price will be a function $$ P(t, T) = A(t, T) e^{-B(t, T) \\cdot r(t)} $$ This is a general form for the analytical solution to the Vasicek model, where $$ B(t, T) = \\frac{1 - e^{-a(T - t)}}{a} $$ $$ A(t, T) = \\exp\\{\\frac{[B(t, T) - T + t](a^2 b - \\tfrac{\\sigma^2}{2})}{a^2} - \\frac{\\sigma^2B^2(t, T)}{4a}\\} $$ Gives you more flexibility, it’s an alternative. There’s one more alternative… ### Cox-Ingersoll-Ross $$ dr = a(b - r) dt + \\sigma \\sqrt{r} dW $$ Again a, b, σ constants. What is the advantage of this model? You can look at the evolution of term structure over time. It’s observed that as the interest rate changes, the volatility also changes. Therefore the model should include this behavior. The price function for CIR has the same general expression as Vasicek: $$ P(t, T) = A(t, T) e^{-B(t, T) r(t)} $$ The functions are just slightly different. We will discuss the similarities and properties of these models. $$ B(t, T) = \\frac{2\\left(e^{\\gamma(T-t)} - 1\\right)}{(\\gamma + a)\\left(e^{\\gamma(T-t)} - 1\\right) + 2\\gamma} $$ $$ A(t, T) = \\left[\\frac{2\\gamma e^{(a + \\gamma)(T - t)/2}}{(\\gamma + a)(e^{\\gamma(T - t)} - 1) + 2\\gamma}\\right]^{2ab/\\sigma^2} $$ where $$ \\gamma = \\sqrt{a^2 + 2 \\gamma^2} $$ Basically, it has an analytical solution. ### Properties of CIR vs Vasicek A(t, T) and B(t, T) are different, but P(t, T) is fundamentally the same general expression. $$ \\frac{\\partial P(t, T)}{r(t)} = -B(t, T) \\cdot P(t, T) $$ This function B can be used therefore as an alternative to duration. From the rate expression, $$ R(t, T) = -\\frac{1}{T - t} \\ln P(t, T) $$ The zero rate at time t for a period T \- t is, if you take the natural log of this general AB expression, because it’s a product, you will get $$ R(t, T) = -\\frac{1}{T-t} \\ln A(t, T) + \\frac{1}{T-t} B(t, T) r(t) $$ Here in this case, the entire term structure can be determined as a function of r(t) *if* a, b, σ are known. Here we define the modified duration of a bond Q. $$ \\hat{D} = B(t, T) $$ The shape is going to be a function of r(t), and dependent on t. The duration of such bond $$ \\frac{\\Delta Q}{Q} = -D \\Delta y $$ If you consider that the yield of a bond. Alternatively, the duration by Vasicek & CIR is $$ \\frac{\\Delta Q}{Q} = - \\hat{D} \\Delta r $$ or $$ \\frac{\\partial Q}{\\partial r} = -\\hat{D} Q $$ ### Example Let’s consider the zero-coupon bond lasting four years. Duration s \= 4, so 10bps parallel shift in the term structure should lead to decrease of 0.4% in the bond price, this is the sensitivity of the model. If we use Vasicek’s model, with a \= 0.1, then $$ \\hat{D} = B(0, 4) = \\frac{1 - e^{-0.1 \\cdot 4}}{0.1} = 3.29 $$ The duration is a function of the mean-reversion rate. That means the short rate is 0.329%. Can you calibrate a using the duration? Hmmmm. Consider Q a portfolio of ZCBs (a coupon-bearing bond). $$ P(t, T_i) (1 \\leq i \\leq m) $$ \\(c_i\\) is the principal of the \\(i^{\\text{th}}\\) bond. Then the duration of our portfolio is $$ \\hat{D} = -\\frac{1}{Q} \\frac{\\partial Q}{\\partial r} = - \\frac{1}{Q} \\sum_{i-1}^m \\frac{\\partial P(t, T_i)}{\\partial r} c_i = \\sum_{i=1}^m \\frac{c_i P(t, T_i)}{Q} \\hat{D}_i $$ Where we have \\(\\hat{D}\\) for a coupon bearing bond being the weighted average of the duration of the underlying ZCBs. ### Vasicek/CIR Pricing Example Suppose we are given a \= 0.1, b \= 0.1. The initial short rate is 10%. The initial stdev of short rate change in a short time Δt is \\(\\sqrt{\\Delta t}\\) Then we will get one point in the term structure, but we can use code to get the entire term structure. ## No-Arbitrage Model These models are designed to be consistent with today’s term structure. Even if it’s used as parameters, it may not be exact. Equilibrium models generate today’s term structure. No-arbitrage models use term structure (the observed rates in the market) as an input. In equilibrium models, drift is not a function of time. In no-arbitrage models, drift *is* a function of time. We will look briefly at the Ho-Lee and Hull-White model. ### Ho-Lee Model $$dr \= \\theta(t) dt \+ \\sigma dW$$ \\(\sigma\\) is the instantaneous standard deviation of the short rate. These parameters are defined so they fit the initial term structure. \\(\theta(t)\\) can be calculated analytically. For the instantaneous forward rate \\(F_t = \\frac{\\partial F}{\\partial t}\\). $$\\theta(t) \= F\_t(0, t) \+ \\sigma^2 t$$ As an approximation, you can take $$\\sigma(t) \\approx F\_t(0, t)$$ This is something kind of similar to what we observe before. The drift is no longer a constant, it is chosen to match the initial term structure. However, the price of a zero coupon bond. $$P(t, T) \= A(t, T) e^{-r(t) (T \- t)}$$ has the analytical solution (skipping derivations) $$\\ln A(t, T) \= \\ln \\frac{P(0, T)}{P(0, t)} \+ (T \- t) F(0, t) \- \\frac{1}{2} \\sigma^2 t (T \- t)^2$$ The advantage of this over previous models is that this will match our term structure. But it’s deficient in some components so it’s not so popular. Assumes that all forward rates have the same stdev ### Hull-White Model $$dr \= \[\\theta(t) \- ar\]dt \+ \\sigma dW$$ or $$dr \= a\[\\tfrac{\\theta(t)}{a} \- r\] dt \+ \\sigma dW$$ a, σ are constant This is the golden goose one factor model. It can be viewed as an extension of the Vasicek model that fits the term structure, with a time-dependent reversion level. Or we can say that it’s similar to Ho-Lee with time-dependent reversion level. $$\\theta(t) \= F\_t(0, t) \+ aF(0, t) \+ \\frac{\\sigma^2}{2a}(1 \- e^{-2at})$$ That last term is small, so we can approximate it. We can say that r follows the slope of the initial instantaneous forward rate curve. Basically this is the partial derivative. Then we can calculate the bond prices, $$P(t, T) \= A(t, T) e^{-B(t, T) r(t)}$$ $$B(t, T) = \\frac{1 - e^{-a(T-t)}}{a}$$ $$\\ln A(t, T) \= \\ln \\frac{P(0, T)}{P(0, t)} \+ B(t, T) F(0, t) \- \\frac{1}{4a^3} \\gamma^2 (e^{-aT} \- e^{-at})^2 (e^{2at} \- 1)$$ ## Some other models Black-Derman-Toy $$d \\ln(r) \= \[\\theta(t) \- a(t) \\ln(r) \] dt \+ \\sigma(t) dW$$ where \\(a(t) = -\\frac{\\sigma'(t)}{\\sigma(t)}\\) Then volatility is related to the speed of mean reversion. Black-Karasinski model is more general, a(t) and σ(t) are determined independently. It has the advantage that interest rates cannot be negative, and the future value is lognormal. However, it’s very difficult to model analytically. These are simple one-factor models as well. The only thing is that they are very difficult to calibrate, especially the more general they get. ### Bond Options We can also have closed form solution for the price of a bond option for Vasicek and Hull-White models: $$LP(0, s)N(h) \- KP(0, T)N(H \- \\sigma\_P)$$ where h and σ are VERY COMPLICATED. The idea is that you have a closed form solution and can solve using analytical methods. ## Assignment 1 (BTW) [https://github.com/pola-rs/polars/issues/5255\#issuecomment-1857124471](https://github.com/pola-rs/polars/issues/5255#issuecomment-1857124471) See this why you can’t have float linspace. Workaround is to multiply by 10 to an integer, floor, cast to int, do an arange, then recast to float and divide by the same magnitude. ## Differential Equations and PDE Methods - URL: https://sharifhsn.dev/blog/computational-methods-week-05/ - Structured data: https://sharifhsn.dev/api/posts/computational-methods-week-05.json - Description: Finite difference methods to approximate PDEs this week - Date: 2025-02-25 - Exact published timestamp: 2025-02-25 - Topics: Computational Methods, Differential Equations, PDEs, Finite Differences, Black-Scholes - Categories: Computational Methods - Source: Computational Methods in Quantitative Finance - Source URL: None ## Plan for the next weeks Finite difference methods to approximate PDEs this week Next week we will continue this and do a little bit more complicated stuff. Homework is due right before spring break. Exam will be any four hours during the weekend, questions will be a little different. ## Differential Equations The first thing is the ODEs, the ordinary differential equations. There is an entire class about this in every program, as long as you do some kind of engineering science you need to take Calc 3, which covers these. The reason they’re called ordinary is because they’re about finding a function \\(f(x)\\). The solution of the equation would be a function, expressed as derivatives of this function. A fake example would be \\(2f''(x) + 3xf'(x) + x^2 + 2 = 0\\) subject to \\(f(0) = 3\\) just to see how it looks. Normally you spend a lot of time solving first-order, the most complicated one. You learn how to solve equations with constant coefficients, and things that can be reduced by transformations to having constant coefficients. Then you make a polynomial that solves the PDE. ## PDEs This is Calc 4, partial differential equations. You would only do this if your area requires it. Not necessarily all of us have done this. I thought this was the hardest, most horrible class. My professor purposely failed people (especially young women) so they would pay him for tutoring. These will usually be a function of more than one variable, like $$ f(t, x) $$ The second variable is usually time, but it doesn’t have to be. x can be multi-dimensional also. Let’s say it’s a line. A practical application would be the heat equation. You have a metal rod and you apply heat to one end, and you look at the distribution of heat over time, f will measure the heat, t will measure time, x will measure location. You could also have a metal plate, and then x is multidimensional. [Embedded figure omitted from the text export.] So \\(f(t, x)\\) is what we’re trying to find. We’re solving in the time \\(t \\in (0, \\infty)\\) Most of the time, these problems have the initial condition. The equations that we will deal with will have a **terminal boundary condition**, ending at T. You know what the option value is at T. Some equations will start at 0, but you can change between boundary condition 0 and T with change of variables. If you do the \\(\\tau = T - t\\), then it will go from terminal to initial. The x we’re talking about is \\(x \\in \\mathbb{R}\\), aka one dimensional, but in general \\(x \\in \\mathbb{R}^n\\). ## NYHOPS Here is a practical example of this. I worked a former Stevens professor some years ago on NYHOPS, New York Harbor Ovserving and Predicting System. Something solves a bunch of PDEs to do forecasts of things like salinity, etc. for the next 48 hours, on a revolving 6 hour period. Actually they just calculate salinity, and it turns out that water speed is driven by salinity. So how does he do it? He uses some other equation that is used for viscosity of water, but we don’t need to worry about it. Because this \\(f(t, x)\\) has two variables, the equations will involve multiple derivatives, **joint derivatives**. LEt’s say we have a **Linear PDE**, which we will define as \\(u(t, x)\\). We will use u for our own purposes. This will only have second order derivatives $$ a \\frac{\\partial^2 u}{\\partial t^2} + b \\frac{\\partial^2 u}{\\partial t \\partial x} + c \\frac{\\partial^2 u}{\\partial x^2} + d\\frac{\\partial u}{\\partial t} + e \\frac{\\partial u}{\\partial x} + fu + j $$ This is the most general form of a linear PDE. Although they don’t have to be constant, for our definitions they need to be constant. The behavior of our solution, it turns out, is only governed by the second order derivatives. So we only need to worry about a, b, and c. We will make a polynomial in α and β, where these represent the derivative with respect to t and x, respectively. Then, $$ P(\\alpha, \\beta) = a\\alpha^2 + b\\alpha \\beta + c \\beta^2 + d\\alpha + e\\beta + ct $$ And again we only worry about a, b, and c. To reiterate: **The nature of the PDE is determined by the properties of these 2nd order terms.** Specifically, we can look at these terms \\(a \\alpha^2 + b \\alpha \\beta + c \\beta^2\\) and divide by \\(\\beta^2\\). THis gives us $$ a(\\frac{\\alpha}{\\beta})^2 + b \\frac{\\alpha}{\\beta} + c $$ And you can see this is a quadratic polynomial. It’s very well-studied and easy to solve. To solve for the determinant, it is $$ \\Delta = b^2 - 4ac $$ We have different behavior based on this Δ. There are three types of major equations. ### Δ \< 0 The equation has no real solutions, only complex conjugate solutions. These lead specifically to **elliptic equations**. One example of this that you would have done if you covered PDEs is the **Laplace equation**. The Laplace equation looks like $$ \\frac{\\partial^2 u}{\\partial t^2} + \\frac{\\partial^2 u}{\\partial x^2} = 0 $$ This is the simplest equation that is elliptic. It appears in thermodynamics. If you’re dealing with gases, behavior in physics, you will see this. ### Δ \> 0 This is called a **hyperbolic equation**. These terms (by the way) come from the behavior of the solution. When you solve a PDE like this, if you don’t have boundary conditions, you get a function which conducts a field. You can look at the derivative of a function and it will tell you where the function goes. You have to tie it down to get one particular surface. The magnetic field is one such function. The minimum is (this is trivial) $$ \\frac{\\partial^2 u}{\\partial t^2} - \\frac{partial^2 u}{\\partial x^2} = 0 $$ And if you pay attention to where a and c are, you can see where the positive and negative come form. This is called the **wave equation**, the jumprope equation which vibrates, useful for earthquakes. The way they do it is by solving an equation of this time, (although with more terms), and the constants are determined by the structure of the earth. ### Δ \= 0 This is called the **parabolic PDE**. The **diffusion equation**, if you’ve heard of diffusion tensor imaging, or something like that, uses these kinds of equations. You can see the image of the brain, but it’s not even correct. The brain is made of grey matter, the neurons, and the white matter, which are the axons that connect them. But nobody can cut open your brain to look at the structure. But the white matter decomposes immediately. Nobody knows how they are connected, so you have to put them in the MRI and see how the water molecules move. These are very intrusive, so you can’t do it for too long, then you have to make connections. You can’t see the cables, you can only see the regions. So someone says to take the PDE from the 1970s where they consider the axons are pipes that they heat up. The **heat equation** is the most classical example: $$ \\frac{\\partial u}{\\partial t} - \\frac{\\partial^2 u}{\\partial x^2} = 0 $$ You have two terms that are missing, that’s the only way to get this equation. If you do this with parameters, then you get a square, which will become transformed into first-order aka not interesting. The only way to get second order you have to keep only one term. And in fact you can flip the t and x here and it’s not a big deal. **All finance of any kind uses this equation**. Because this is basically what you get when you have Markov processes. And in general diffusion of particles involves this. ## Boundary Conditions In general, PDEs are solved for \\(t \\in (0, \\infty) \\times x \\in \\mathbb{R}\\). So how does this space look? [Embedded figure omitted from the text export.] But in finance, we are bounded by T: If I don’t specify the boundary condition, the solution is floating, it can be infinite number of curves. I have to tie it down with the boundary T. And there’s another boundary at x \= 0\. You tie it down from three different places, t \> 0, t \< T, and x \> 0\. And I don’t care about anything other than when t \= 0, because that’s where I am now. [Embedded figure omitted from the text export.] In a nutshell, to solve this, we create a domain, and we make a grid on this domain. In order for this grid to fit, we have to make squares. They solve the equation on each tiny square. Then they propagate the solution (I will show you how soon). This is nothing special, it is general and it applies to any PDE whatsoever. ## Methodology First: what is the PDE? Let’s call \\(V(t, S)\\) for the value of an option at time t with asset price S. We know that if the asset follows risk-neutral geometric Brownian motion, then V solves the following PDE: $$ \\frac{\\partial V}{\\partial t} + rS\\frac{\\partial V}{\\partial S} + \\frac{1}{2} \\sigma^2 S^2 \\frac{\\partial^2 V}{\\partial S^2} - rV = 0 $$ And this is the Black-Scholes PDE (technically, it’s Merton’s because he’s the one that did the PDE. However this does not look like the heat equation, because the coefficients are not constant. But you can make it constant through the use of transformation: \\(S = e^x\\), which is the same as \\(x = \\ln S\\) You can make the t go to T to solve the initial value as well, but this is a different way. Then we can call $$ V(S, t) = V(e^x, t) = u(x, t) $$ $$ \\frac{\\partial V}{\\partial t}(t, s) = \\frac{\\partial u}{\\partial t}(t, x) $$ Since t doesn’t do anything, this is easy. But what is \\(\\frac{\\partial V}{\\partial S}\\) in terms of x? I have to substitute two derivatives. But you can do the chain rule to solve for dvdS $$ \\frac{\\partial u}{\\partial t} + (r - \\frac{\\sigma^2}{2}) \\frac{\\partial u}{\\partial x} + \\frac{1}{2} \\sigma^2 \\frac{\\partial^2 u}{\\partial x^2} - ru = 0 $$ And we’ll call r \- σ^2/2 \= μ for convenience. Because this is constant coefficients, you can actually make this into a heat equation which is easy to solve. Merton solves it in a horrible way which is not like that. And you can reduce it into Black-Scholes. But if there is no analytical solution, how do I approximate this solution? For Black-Scholes, we already have the analytical solution. We learn it as an example here, and once we learn the principle we can apply the methods for problems where there are only numerical solutions. NYHOPS does not have an analytical solution. There’s two different ways of solving this. There is the **explicit** and **implicit** way. They are both **finite difference methods**. There is another way which is more appropriate to NYHOPS, but we will not cover that. ## Explicit Finite Difference First, we take our domain. $$ t \\in [0, T) $$ x is different because we did this logarithm. $$ x = \\ln S = (-\\infty, \\infty) $$ We need to discretize this domain. A square grid is the easiest, simplest way to do it. Most PDEs are solved in this way. We substitute the derivative with the finite difference at each point on the grid. Then we can solve based on where the asset price is. In the explicit method, we use the boundary, and you move from there to the direction you want to go. And you do so explicitly (Will explain soon what that means). Depending on how many terms you have, you need t terms to get one term, in this case 3\. [Embedded figure omitted from the text export.] ## Implicit The difference in this equation, is that you won’t be able to explain each point in terms of one. You have to solve equations in general. Explicit describes the points explicitly in terms of previous points. For implicit, the value is expressed by the value implied by ALL the values on the previous points. ## Differences Generally, explicit is easier, but it may not converge. Implicit always converges. BTW, NYHOPS works with curves and boundaries, which is called the finite element method. When you propagate these, you get inconsistencies, so if you have to keep propagating until you get convergence. You solve for the vertices of the cube. Now you have nine points available. ## How This is the actual math for doing this. ### Discretize the Domain We need to discretize t and x. $$ \\Delta t = \\frac{T}{n} $$ very simple. Δx can be anything, with smaller being better. Is there an ideal relationship between Δx and Δt? There is, but we worry about that later. $$ t = (0, \\Delta t, 2 \\Delta t, \\ldots, n \\Delta t) $$ $$ x = (-N \\Delta x, (-N + 1) \\Delta X, \\ldots, 0, \\Delta x, \\ldots, N \\Delta x) $$ So we have 2N \+ 1 points in x, and n \+ 1 points in t. t thing doesn’t matter, x is crucial. And we will notate the value of the function u as $$ t_i = i\\Delta t $$ $$ x_i = j\\Delta x - N $$ $$ u(t_i, x_j) = u_{ij} $$ We need three derivatives ### Explicit This one is easy. The big difference is in how I make my derivatives. They are calculated directly from these three points: [Embedded figure omitted from the text export.] Here are our equations: $$ \\frac{\\partial u}{\\partial t} = \\frac{u_{i+1,j}-u_{ij}}{\\Delta t} $$ $$ \\frac{\\partial u}{\\partial x} = \\frac{u_{i+1, j+1} - u_{i+1, j-1}}{2\\Delta x} $$ $$ \\frac{\\partial^2 u}{\\partial x^2} = \\frac{u_{i+1,j+1}-2u_{i+1,j}+u_{i+1,j-1}}{\\Delta x^2} $$ To get it, you need those three points. All of these things are now substituted in our PDE. $$ \\frac{\\partial u}{\\partial t} + \\mu \\frac{\\partial u}{\\partial x} + \\frac{1}{2} \\sigma^2 \\frac{\\partial^2 u}{\\partial x^2} - ru = 0 $$ Just the points with some constants like μ and σ. Now you can determine those constants in terms of the other terms: $$ u_{ij} = \\Delta t(\\frac{\\sigma^2}{2\\Delta x^2} + \\frac{\\mu}{2\\Delta x}) \\mu_{i+1,j+1} + (\\text{a term})u_{i+1,j} + \\text{another term}u_{i+1,j-1} $$ If we take those coefficient terms to be \\(P_u\\), \\(P_m\\), and \\(P_d\\), respectively, the probability of going up, staying the same, or going down, then it turns into a trinomial. And it’s almost the same as the formula for the trinomial tree\! The only difference is that explicit finite difference has \\(\\frac{1}{1+r\\Delta t}\\) inside the formulas, whereas the trinomial uses \\(\\frac{1}{e^{r\\Delta t}}\\) aka discrete vs continuous. And of course as N increases we approach continuous. ## Stability and Convergence In order for this thing to converge, we must have \\(\\Delta x \\geq \\sigma \\sqrt{3\\Delta t}\\) which is the same condition as the trinomial tree. You can make Δx small, but not too small. Conceptually, Δx tells me how many points I have to solve for, and Δt tells me the number of steps I have to go through the tree. Because we have 3 to 1, we must have N \> n. We’re working around the grid, and N tells me how many points I have. And you can see this visually, that you will get stuck if you have too many Δx compared to Δt: [Embedded figure omitted from the text export.] If I pick \\(\\Delta x = \\sigma \\sqrt{3 \\Delta t}\\). If you calculate Δ you need two points in the origin, Γ you need three points. Sometimes that is useful. ## Implicit Scheme Fundamentally, if you understand the explicit, it’s similar. Keep in mind the points. We’ll be using FOUR POINTS for the three derivatives. [Embedded figure omitted from the text export.] It’s hard to remember formulas. But it’s easier to remember how to derive them. Now I’m going to plug them into the PDE, and then solve it. Remember that these coefficients are constants. Then you will get an equation relating these four points $$ Au_{i,j+1} + Bu_{i,j} + Cu_{i,j-1} = u_{i+1,j} $$ The NUMBER ONE MOST IMPORTANT THING IS: **A, B, C are the same for all i,j**. This is normal because the coefficients of the PDE don’t depend on the location or time, no t or x in it. That makes it simpler to solve\! Nonetheless, this is still an implicit solution. **A, B, C are not probabilities anymore**. We previously generalized American options by calculating the expected value of the future value of the option given that you are at that point. So you could take that point and compare what happens at exercise and store it. However, this doesn’t work, so it loses the interpretation of expected value. **There will be an exercise about this**. We have three unknowns and one equation, so we have to solve a lot of equations. We have 2N \+ 1 points. Because we are writing one equation for each set of three points, we will be lacking two equations, so we have 2N \- 2 equations. So the system built here CANNOT be solved. Therefore we need to come up with two more equations. The two extra equations are coming from boundary conditions. Remember, N is supposed to be very large. And our underlying is the log of the stock, these boundaries relate to very high and very low stock value. It depends, because for a call, when a stock rises in value, the call becomes very valuable, for a put, it becomes worthless. For Put, when \\(S \\uparrow \\infty\\), then value is small. Under our property of log, it can’t be 0, although obviously we think it is. $$ \\frac{\\partial V}{\\partial S} = 0 $$ because it’s not going to change at all. And then when \\(S \\downarrow 0\\) $$ \\frac{\\partial V}{\\partial S} = -1 $$ This comes from the fact that K is a constant and derivative of S by itself is 1\. If we consider the topmost \\(i,N\\). The difference $$ \\frac{u_{i,N} - u_{i,N-1}}{e^{N\\Delta x + x_0} - e^{(N-1)\\Delta x + x_0}} = 0 $$ We should not be using u, but the corresponding Vs. We need to use the unknowns that we have, not other unknowns. So we can’t use dV/dS, we have to use the corresponding du/dS. And we can get an extra equation from $$ u_{i,N} - u_{i,N-1} = 0 $$ For the bottom put, it’s $$ \\frac{u_{i,-N} - u_{i, -N+1}}{e^{(-N+1)\\Delta x + x_0} - e^{-N\\Delta x + x_0}} $$ and that becomes $$ u_{i, -N} - u_{i, -N+1} = \\lambda_D $$ where λ is some number. $$ u_{i,N} - u_{i,N-1} = \\lambda_U $$ $$ u_{i, -N} - u_{i, -N+1} = \\lambda_D $$ And these are two extra equations that we can use to solve. But this matrix is HUGE, 2N+1 by 2N+1, will cripple your computer. We deal with imbeciles that solve by exhaustive search, that’s how LLMs work. If you take these matrix 201x201 and ask R to solve it, it’s very fast. Then 1000x1000 is a little slower, but it can still do it. However, you can solve this with your brain, by derivation. It is three formulas, which are easy to implement\! Since basically only the diagonal is populated, we are wasting our time doing typical matrix solving. The idea is kinda cool. We know the matrix is solvable, because it’s invertible. We know that because the determinant exists and is not 0, easy to show. There is a general method called the **Jacobian**, where you come up with some random numbers and plug them in. But there are two other methods to solve, the 5th grade ones: substitution and elimination. And you can solve it in this way. The first equation is $$ u_N u_{N-1} $$ You can express the latter in terms of the former.. Then you can express \\(u_{N-2}\\) in the same way, and keep bootstrapping to make it all functions of \\(u_N\\). Then you go to the end and get \\(u_{N-1}\\) and \\(u_{-N}\\) and get \\(u_N\\) from there, and then get everything else from there. This is called in the book **solving a tridiagonal system**. Because you have three main diagonals, this is possible. Technically you can solve this in general. And this is in the book too. We will assume matrix A [Embedded figure omitted from the text export.] You’ll see that this works with the numbers being all different, although it’s easier in our case with the numbers being the same. To solve this, we go through the motion: $$ a_{11} x_1 + a_{12} x_2 = y_1 $$ $$ x_1 = \\frac{1}{a_{11}} y_1 - \\frac{a_{12}}{a_{11}}x_2 $$ Let’s look at this structure. We have a number \- another number times x\_2. We’ll call those numbers C\_1 and D\_1. $$ x_1 = C_1 + D_1 x_2 $$ Then $$ x_2 = \\frac{y_2 - a_{21} C_1}{a_{21}D_1 + a_{22}} - \\frac{a_{23}}{a_{21} D_1 + a_{22}} x_3 $$ And you can see it’s the same type of expression as x\_1. I came up with this on my own, the general idea of the derivation. In general, each step i follows $$ x_i = C_i + D_i x_{i+1} $$ where $$ C_i = \\frac{y_i - a_{i, i -1} C_{i-1}}{a_{i, i-1} D_{i - 1} + a_{i,i}} $$ and $$ D_i = -\\frac{a_{i, i+1}}{a_{i, i-1} D_{i-1} + a_{ii}} $$ In general, \\(y_i\\) is $$ a_{i+1, i}x_i + a_{i+1,i+1} x_{i+1} + a_{i+1,i+2} x_{i+2} = y_{i+1} $$ we can substitute the same C and D terms as we had done in 2, and turns out to be the exact same terms. That means it’s really easy to program\! The only thing is, the end gets treated a little differently. The last equation: $$ a_{n, n-1} x_{n-1} + a_{nn} x_n = y_n $$ We take that x value, and we say it’s equal to $$ x_{n-1} = C_{n-1} + D_{n-1} x_n $$ It makes it easy to express in the future, because then you can move them all to the end. Once you plug that in backwards, you’re going to get x\_n, so $$ x_n = \\frac{y_n - a_{n,n-1} C_{n-1}}{a_{n,n-1} D_{n-1} + a_{nn}} $$ AND THEN I’M DONE\! ## Black–Scholes Pricing for FX - URL: https://sharifhsn.dev/blog/fx-black-scholes-pricing/ - Structured data: https://sharifhsn.dev/api/posts/fx-black-scholes-pricing.json - Description: The FE-635 workbook implements Black–Scholes-style pricing for calls, puts, forwards, and deposits. Its inputs separate the domestic discount rate from the foreign rate, because an… - Date: 2025-02-24 - Exact published timestamp: 2025-02-24 - Topics: FX, Black–Scholes, Option Pricing - Categories: FX - Source: FE-635 \| Risk Engineering - Source URL: None The FE-635 workbook implements Black–Scholes-style pricing for calls, puts, forwards, and deposits. Its inputs separate the domestic discount rate from the foreign rate, because an FX option has two money-market accounts in the carry relationship. The spreadsheet's maturity convention is explicit: \(T=\text{Days}/365\). The pricing routine then consumes forward or strike, domestic and foreign rates, and volatility. That separation is more important than the function name; passing spot where the workbook expects forward FX changes the result systematically. The notes treat the implementation as a practitioner tool. It is a compact expression of the assumptions, not a substitute for checking quote direction, settlement, and discounting currency. ## Interest Rate Adjustments and Currency Swaps - URL: https://sharifhsn.dev/blog/advanced-derivatives-week-04/ - Structured data: https://sharifhsn.dev/api/posts/advanced-derivatives-week-04.json - Description: We may have to make some adjustments when we value interest rate derivatives. These are convexity, timing, and quant adjustments. We will see when it is appropriate to make these a… - Date: 2025-02-20 - Exact published timestamp: 2025-02-20 - Topics: Fixed Income, Interest Rate Derivatives, Convexity Adjustment, Timing Adjustment, Change of Numeraire, Currency Swaps - Categories: Fixed Income - Source: Advanced Derivatives - Source URL: None ## Adjustments We may have to make some adjustments when we value interest rate derivatives. These are convexity, timing, and quant adjustments. We will see when it is appropriate to make these adjustments. ## A general two-step procedure for valuing a European style derivative First, calculate the expected payoff by assuming that the expected value of each underlying variable equals its forward value. Then, we discount it based on the interest rate. For non-standard interest rate derivatives, we need to modify the first step with adjustments to the forward value If you look at the forward yields and forward prices, this kind of convexity adjustment may be necessary to make when the derivative is dependent on the yield of the bond. Basically, we define the forward yield on a bond as being the yield calculated from the forward bond price. There is a non-linear relationship between bond yields and bond prices. It follows that when the forward bond price equals the expected future bond price, the forward yield does not necessarily equal the expected future yield, you cannot have both. Suppose: B\_T \= price of a bond at T y\_T \= yield of a bond at T We have a relationship here that we can write as B\_T \= G(y\_T) where G is a nonlinear function F\_0 \= forward bond price at time 0 for a transaction maturing at time T y\_0 \= forward yield at time 0 We have a relationship where F\_0 \= G(y\_0) ![Hand-drawn bond price and bond yield curves from the source notes.](/static/img/Advanced Derivatives-week-04-bond-price-yield.png) If we make our forward equivalent to the expected future bond price, then $$F\_T \= \\mathbb{E}\[B\_T\] \= B\_2$$ Then our yield is $$\\mathbb{E}\[y\_T\] \= \\frac{1}{3} \\sum\_{i=1}^3 y\_i$$ Such value is going to be greater than y\_2. ## Convexity Adjustment Since we are incorporating convexity, we will adjust the expectation of the yield Suppose that the payoff from a derivative at time T depends on the bond yield observed at T. Define y\_0 \= forward bond yield for a contract maturing at T y\_T \= bond yield at T B\_T \= price of bond at T σ\_y \= volatility of forward bond yield We have the relationship previously expressed that B\_T \= G(y\_T) One thing we can do is expand G(y\_T) in a Taylor series about y\_0. Here we can write \\(B_T \\approx G(y_0) + (y_T - y_0)G'(y_0) + \\tfrac{1}{2}(y_T-y_0)^2 G''(y_0)\\) I can take the expectation of both expressions, left and right, in a world that is forward risk neutral with respect to a zero-coupon bond, maturing at T. $$\\mathbb{E}_T[B_T] = G(y_0) + \\mathbb{E}_T[y_T - y_0]G'(y_0) + \\frac{1}{2}\\mathbb{E}_T[(y_T-y_0)^2]G''(y_0)$$ What is G(y\_0)? F\_0, the forward bond price. We would say the forward is the expected value. But this assumption does not work for the yield. Let us assume anyway. F\_0 \= E\_T\[B\_T\] Now let’s crack this apart. We can see that y\_T \- y\_0 is a constant. Also, E\_T\[(y\_T \- y\_0)^2\] \= σ\_y^2 y\_0^2 T. Not sure why… maybe something to do with quadratic variation? I will post the proof online, using Itô’s lemma. $$\\mathbb{E}_T[y_T] = y_0 - \\frac{1}{2}y_0^2 \\sigma_y^2 T \\frac{G''(y_0)}{G'(y_0)}$$ This ½ term is considered the “convexity adjustment”. It’s the difference between the expected bond yield and the forward bond yield at t=0. Let’s assume that we have a GBM, and that we have a numeraire at time T. Suppose that the growth rate of forward bond yield is denoted by α and its volatility $$\\sigma\_y$$. For example, we can use Itô’s lemma to calculate the process for the forward bond price. $$dy \= \\alpha y dt \+ \\sigma\_y y dW$$ From Itô’s lemma we get $$ d[G(y)] = [G'(y)\\,dy + \\frac{1}{2}G''(y)\\sigma_y^2y^2]dt + G'(y)\\sigma_y y\\,dW $$ Given that the expected growth rate of G(y) is zero, We can find that $$G'(y) \\alpha y + \\frac{1}{2} G''(y) \\sigma_y^2 y^2 = 0$$ or $$\\alpha = -\\frac{1}{2}\\frac{G''(y)}{G'(y)} \\sigma_y^2 y$$ which results in our final convexity adjustment. ## Applications for Interest Rate Derivatives Consider an instrument that provides a cash flow at time T equal to the interest rate between T and t\* applied to principal L. The cash flow at time T is going to be equal to the principal and rate and tau $$LR\_T\\tau$$ where τ is T\* \- T. In order to apply this convexity adjustment, we have the relationship between price and yield. For a ZCB, G(y) is $$G(y) \= \\frac{1}{1+y\\tau}$$ From the previous equations, the expected value of such a rate is $$\\mathbb{E}_T[R_T] = R_0 - \\frac{1}{2}R_0 \\sigma_R^2 T\\frac{G''(R_0)}{G'(R_0)}$$ or, if you take the derivative of such expectations, you get $$\\mathbb{E}_T[R_T] = R_0 + \\frac{R_0 \\sigma_R^2 \\tau T}{1 + R_0 \\tau}$$ where $$R\_0$$ is the forward rate applicable between T and T\*. Then the present value of the instrument is $$P(0, T) L \\tau \[R\_0 \+ \\frac{R\_0 \\sigma\_R^2 \\tau T}{1 \+ R\_0 \\tau}\]$$ ## Numeric Example of Convexity Adjustment An instrument: payoff in 3 years equal to 1 year zero coupon rate multiplied by $1000 vol is 20%, yield curve is flat at 10%, annual compounding, convexity adjustment is 10.9 bps. Value of instrument is 75.95 If we have T \= 3, T\* \= 4, τ \= 1, R\_0 \= 0.10, σ\_R \= 0.20 Value of derivative is $$1000 \\times 1 \\times \[0.10 \+ \\frac{0.10^2 \\cdot 0.20^2 \\cdot 1 \\cdot 3}{1 \+ 0.10 \- 1}\]$$ And you also have to discount the payoff, the yield curve is flat so it ends up being $$\\frac{1}{1.10^3}$$ ## Convexity Adjustment for Swap Rate The adjustment is for the life of the swap. The derivative is on the swap observed at T, the life is from T to T \+ τ. What is different here? Just the reference instrument. A direct application Swap rate is approximated as the yield on a 12% bond. $$G(y) \= \\frac{0.12}{(1+y)} \+ \\frac{0.12}{(1+y)^2} \+ \\frac{1.12}{(1+y)^3}$$ $$G'(y) = -\\frac{0.12}{(1+y)^2} - \\frac{0.24}{(1+y)^3} - \\frac{3.36}{(1+y)^4}$$ Furthermore, $$G''(y) = \\frac{0.24}{(1+y)^3} + \\frac{0.72}{(1+y)^4} + \\frac{13.44}{(1+y)^5}$$ In this case, the forward yield y\_0 \= 0.12. G'(y_0) = -2.4018; G''(y_0) = 8.2546. Then plug it into the convexity adjustment equation: will get you 0.1236, 12.36% Therefore the value of the instrument is The basic procedure is to find the nonlinear function G, then take its derivatives, then plug the parameters in to this equation. ## Timing Adjustment Consider a case where a market variable **v** is observed at time T and its value is used to calculate the payoff that occurs later at time T\*. Define v\_T \= value of v at time T $$\\mathbb{E}\_T\[v\_T\]$$ \= the expected value of v\_T in a world that is forward risk neutral (FRN) with respect to P(t, T) $$\\mathbb{E}_{T\*}[v_T]$$ = the expected value of v_T* in a world that is forward risk neutral (FRN) with respect to P(t, T\*) $$w \= \\frac{P(t, T\*)}{P(t, T)}$$ the ratio of the ZCB at different times, equal to the forward price of ZCB from T to T\*. σ\_v \= volatility of v σ\_w \= volatility of w ρ\_vw \= correlation between v and w So then what is this expectation at T\*? $$\\mathbb{E}\_{T\*}\[v\_T\]$$ ## Change of Numeraire One way to look at this is by discussing the change of numeraire, what is the impact of that, if we’re following the market variable? Assume that the variable is the price of a traded security **f** in a world where the market price of risk is λ\_i. $$df \= \[r \+ \\sum\_{i=1}^n \\lambda\_i \\sigma\_{f,i} \]fdt \+ \\sum\_{i=1}^n \\sigma\_{f, i} dW\_i$$ When market price of risk is λ\_i\* $$df \= \[r \+ \\sum\_{i=1}^n \\lambda\_i\* \\sigma\_{f,i} \]fdt \+ \\sum\_{i=1}^n \\sigma\_{f, i} dW\_i$$ Then what is the effect of moving from the first world to the second world (star world)? The expected growth rate of price of any traded security f. $$\\sum\_{i=1}^n (\\lambda\_i\* \- \\lambda\_i) \\sigma\_{f, i}$$ Define $$w \= \\frac{h}{g}$$ ## Quantos Derivatives where the payoff is defined using variables measured in one currency, and paid in another currency. Essentially this represents the exchange rate. It’s very similar to the timing adjustment. No adjustment is necessary for vanilla swap, cap, or swaption. ## Skipped Nonstandard, rest of stuff… ## Currency Swaps A swap for LIBOR in two currencies is theoretically worth 0\. However, this will have a spread in the exchange typically. So we need to adjust for this. There are also more complex swaps. They cannot be valued by assuming that forward rates will be realized. LIBOR-in-arrears swap, constant maturity swaps (CMS/CMT) ### LIBOR-in-arrears Rate is observed and paid at time T, not T\*. Convexity adjustment to each forward rate underlying the swap. This one is dependent on time when you have resets. ![Handwritten sketch from the Week 4 source notes.](/static/img/Advanced Derivatives-week-04-rate-adjustment.png) ## PBMs Under Fire: What CVS's Response Leaves Out - URL: https://sharifhsn.dev/blog/linkedin-2025-02-18-pbms-under-fire-what-cvs-s-response-leaves-out/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-02-18-pbms-under-fire-what-cvs-s-response-leaves-out.json - Description: Trump has taken aim at pharmacy benefit managers, and the response from the industry seems to be… duck and cover. Prem Shah, executive VP at CVS, responded to skepticism on #Bloomb… - Date: 2025-02-18 - Exact published timestamp: 2025-02-18T16:47:24.726Z - Topics: Healthcare, Public Policy - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7297657973143666688/ Trump has taken aim at pharmacy benefit managers, and the response from the industry seems to be… duck and cover. Prem Shah, executive VP at CVS, responded to skepticism on #Bloomberg today. Pharmacy benefit managers (PBMs) act as a middleman between purchasers of prescription medicine, like health insurers. They negotiate directly with drug manufacturers and maintain drug formularies (the list of medicines covered by insurance). Theoretically, they’re able to consolidate these services for a large number of purchasers and save money that way. That’s certainly what Prem Shah says. He repeated the claim found on the industry website that PBMs are responsible for over $1000 in savings per person per year. [1] But that is decidedly at odds with the recent FTC report on PBMs, which found that they marked up cancer and HIV-related drugs by thousands of percent, resulting in $7.3 billion in additional revenue for the “Big 3” (CVS, ESI, and OptumRx). A previous report found that the percentage of dispensing revenue for specialty drugs like these going to PBMs increased from 54% to 68% between 2016 and 2023. [2] Shah says that he hopes he can have a dialogue with the Trump administration regarding the value that PBMs provide. Time will tell how successful his efforts will be. What do you think? Should these middlemen continue to exist? [1] [https://lnkd.in/eUfh9PPN](https://lnkd.in/eUfh9PPN) [2] [https://lnkd.in/eyNTjfq7](https://lnkd.in/eyNTjfq7) ## Trinomial Trees, Dividends, and Greeks - URL: https://sharifhsn.dev/blog/computational-methods-week-04/ - Structured data: https://sharifhsn.dev/api/posts/computational-methods-week-04.json - Description: I posted an assignment, this assignment is due this week on Sunday. Let’s just talk about the homework. - Date: 2025-02-18 - Exact published timestamp: 2025-02-18 - Topics: Computational Methods, Trinomial Trees, Dividends, Greeks, Finite Differences - Categories: Computational Methods - Source: Computational Methods in Quantitative Finance - Source URL: None ## Homework I posted an assignment, this assignment is due this week on Sunday. Let’s just talk about the homework. This covers material from first two lectures. This lecture and previous lecture (trees) will be HW 2\. Q: “Implement the Newton \-whatever method, do we have to implement them alL?” A: You can implement all three for bonus Q: Put, vs call option? A: Use both Q: When you see equity data in the first part, what should we get? Tick-level data, just the price? A: You need price of the underlying. You don’t need for daily or anything, the most important option derivative is delta and gamma. Basically, you can calculate the price of the option if you know the price of the stock. You really needed to read the price of the stock, just get that price at the moment you download the option data. Q: Do we have to include code in the appendix? A: You don’t need an appendix, but if you use Python, you should include the py file or Jupyter notebook, some way of checking that? Q: How many strike prices do we need? A: You should use whatever makes sense to you. Normally, let’s use VIX as an example. That’s calculated from options on S\&P 500\. They literally take all of the options, because there’s tons of them, until they get to the two that have had zero trades in the last five minutes or so. That’s the rules. They basically get to options that haven’t been traded in a long time, so their prices are not current. Look at the most traded options, and don’t include options that don’t have a lot of volume. Just go enough in the tails to make sense for the data. What I want you to see in the option data and detail in the submission is the volatility smirk. I want to see some kind of decrease with strike price, and what happens when you look at next month volatility vs three months, then plot them all and see what you observe. Then it’s possible that the day you downloaded the data you don’t observe that, and that’s okay\! The point is to explain what you get from the data. ## Forgot to Discuss Last Time When we calculate this option data, we said well, you basically step down in the tree, and then at every step you calculate the discounted expected value based on the nodes further on in the tree. \[In Romania we have to do a thesis after Bachelor’s degree. In 1995, at the time, something just appeared. Before you type with a typewriter or write nicely with a pen. Computers had appeared, so Windows 3.1 is what we were using. It had a Word document processor. I was saving the files on the floppy disk. I was using the school computer. After I wrote half of my thesis, the file just crashed. So I couldn’t really open it. I was writing it with a pen, and I was typing it, so it was horrible, and I was about to quit school. And then a friend of his said that he was working at some company and said “I don’t know anything”. He was in charge of the Bucharest water system and routing teams to fix water when it breaks. You may not realize, but there’s a lot of emergencies that happen in a big city. He said that because I’m not doing anything else, I’ll type your thesis. And because of this experience, I have to save things. You guys are spoiled, everything is saved online\] ## Combinatorial Formula {#combinatorial-formula} There’s not really a big deal about this. We discussed trees and how they are constructed. The tree represents the volution of the stock. [Embedded diagram omitted from the text export.] You have to go down at the first step, then go all the way up. Or go down at the second step, then go up the rest. Or go down at third step after happening. You can think about it in a combinatorial way, where it’s $$(3 \\choose 1)$$ All of these paths have the same probability. The probability that you end up in this particular node over here will be $$(3 \\choose 1\) \* p\_u^2 \* p\_d^1$$ And this is exactly the binomial probability. In general, if you look at n steps, at the top value, you know what the value is. Let’s say it’s an additive tree. [Embedded diagram omitted from the text export.] Because of this, you can calculate the payoff just by using the top price $$\\varphi(e^{x\_0 \+ n \\Delta x\_u}) \\times (n \\choose 0\) p\_u^n$$ This is the probability of the one node. My payoff is $$\\varphi(S\_T)$$, but it’s $$\\mathbb{E}\[\\varphi(S\_T)|\\mathcal{F}\_0\]$$ Conditioned on time 0\. So my expectation is payoff times probability. $$= \\sum\_{i=0}^n \\varphi(e^{x\_0 \+ i \\Delta x\_u}) \\times (n \\choose i) p\_u^{i} p\_d^{n-i}$$ You are creating discrete paths for your process, which matches, in the limit, as n approaches infinity, the path of the general process. That’s the combinatorics formula. This formula doesn’t work if you price American options because they’re path-dependent. So I don’t know if I will end up with that final payoff, we have to calculate every step of the way. ## Greeks ### δ $$δ \= \\frac{\\partial C}{\\partial S}$$ The partial derivative of the call option with respect to the stock price, its sensitivity. You’re supposed to calculate the change in the option price when the current price changes. What most of the interviews will say, is to do the tree. [Embedded diagram omitted from the text export.] You would take the distance, the difference in actual values. When you step back in the tree, you get the call value, and you will actually calculate all these values. You can then approximate this Delta as $$\\frac{c(S\_0 u) \- c(S\_0 d)}{S\_0 u \- S\_0 d}$$ This is not a very good approximation because you’re not calculating the delta now, you’re calculating it in the future. In the interview, you will construct the three step tree, and you’ll have the number and differences, and very easy to calculate. This is not correct though. It’s supposed to be some value with very little difference in the stock. So you should calculate 2 trees\! One starts that from S\_0, and one that starts from S\_0 \+ ΔS, and S\_0 \+ ΔS. This will allow you to calculate the first and second derivative. You need the third point for the gamma (second derivative). Understand the idea\! Don’t just apply what you read. Why am I doing it this way? There is no mystery to this, like math, it has to be logical. If it’s not logical, then it’s wrong. This is a method that is very useful if you write it on a piece of paper. If you have a computer with you, a call using a tree takes a fraction of a second. So why wouldn’t I do two in a fraction of a second? It’s pretty easy to do the actual calculation, but this was not possible many years ago. All the other derivatives are the same idea. All you do is vary the underlying derivative slightly, and hold everything constant. Some are them are irrelevant, but you can\! That’s the one thing to mention from last class. ## Trinomial Tree It’s just the same thing as the binomial tree. There is literally no difference. First of all, I want you to remember how we did the binomial tree. $$\\mathbb{E}\[\\Delta R^{\\text{disc}}\] \= \\mathbb{E}\[\\Delta R^{\\text{cont}}\]$$ And same thing for variance. And we did some other stuff to make sure that the probability distribution is correct. So we will use the same exact idea. Remember that from the continuous process dR\_t is the logarithm of the stock price. $$dR\_t \= (r \- \\frac{\\sigma^2}{2}) dt \+ \\sigma dW\_t$$ $$\\mathbb{E}\[\\Delta R\_t\] \= (r \- \\frac{\\sigma^2}{2}) \\Delta t$$ $$\\mathbb{V}\[\\Delta R\_t\] \= \\sigma^2 \\Delta t$$ $$\\mathbb{E}\[\\Delta R\_t^2\] \= \\sigma^2 \\Delta t \+ (r \- \\frac{\\sigma^2}{2})^2 \\Delta t^2$$ How do we do this? Two conditions: - Has to converge - Has to be recombining Same number of nodes, two above, and two below. 2n \+ 1 is still a linear increase. [Embedded diagram omitted from the text export.] That would be convenient. If I did like this [Embedded diagram omitted from the text export.] It would be O(n^2), really bad. For the recombining one, it would be It turns out the condition you need to have is $$\\Delta R\_u \+ \\Delta R\_d \= 2\\Delta R\_m$$ You can solve the whole thing with this. The problem is that you’re going to get a huge system. Six unknowns, and three equations. No matter what, we have three equations. One with the expectation, one with the variance, and probabilities must equal to 1\. The six unknowns is the three ps and three Rs. This will have a huge number of solutions that will give us a trinomial tree. We won’t do that. We’re going to say that this is too complicate,d so we will make on etree. We will take ΔR\_u \= ΔR, ΔR\_d \= \-ΔR, ΔR\_m \= 0\. [Embedded diagram omitted from the text export.] Now we have four unknowns, and three equations, so still undetermined. But what’s going to happen is we will write the equations. ### Calhoun Sidebar some ex-students want to form a business, get help from students. $$\\Delta R p\_u \- \\Delta R p\_d \= (r \- \\frac{\\sigma^2}{2}) \\Delta t$$ And same thing for $$\\Delta R^2 p\_u \+ \\Delta R^2 p\_d \= \\sigma^2 \\Delta t \+ (r \- \\frac{\\sigma^2}{2})^2 \\Delta t^2$$ (here is supposed to be \-ΔR, but squared is he same thing) $$p\_u \+ p\_m \+ p\_d \= 1$$ I have a choice of how large my tree should step. Technically, you can solve for everything(not really) If you look at the probabilities, you can calculate expected value in terms of the parameters. Once you have the formula, $$p\_u \= \\frac{1}{2} (\\frac{\\sigma^2\\Delta t^2 \+ (r \- \\tfrac{\\sigma^2}{2})^2 \\Delta t}{\\Delta R^2} \+ \\frac{(r \- \\tfrac{\\sigma^2}{2}) \\Delta t}{\\Delta R})$$ This probability has to be a number between 0 and 1\. And you can guarantee this with a certain condition. In order for this to be true, we need a sufficient condition that $$\\Delta R \\geq \\sigma \\sqrt{3\\Delta t}$$ This is not easy to get, and it is sufficient, not necessary, so you can get tighter guarantees. It’s important to understand what’s happening here fundamentally. [Embedded diagram omitted from the text export.] The more steps I add, that’s what Δt is. But I can’t make the increments too small, it must have some kind of spread. When you’re guaranteeing a minimum of spread, it captures the distribution, without that you don’t get the distribution. This formula is calculated numerically, but this is the intuition. ## Order of Convergence Order of convergence of the trinomial tree is O(ΔR^2 \+ Δt). What does that mean? As Δt goes to 0, ΔR also has to go to 0 at the same rate. If I take Δt as 0.01, then I need ΔR to be greater than 0.1 based on that formula above. If I have ΔR \= 0.5, to square it 0.25, it will get too close. So I actually have to take them in the same convergence. And the best ΔR uses this formula, so it’s just EQUAL to that square root. This makes sure that the tree has the fastest convergence. Now the tree is unique, you calculate ΔR, you get the probabilities, and it’s a unique tree that satisfies these assumptions. Professor, you didn’t explain to us what big O is. This means that the call option value from the tree, minus the true value (unknown), the absolute value of that should be approximately C(ΔR^2 \+ Δt), where C is a constant. The constant screws me up, nobody knows what it is. Even if you make your tree within one cents, if there’s a large constant it’s worthless. You can’t get exact numbers. This is for the trinomial tree. If we do this same thing for binomial, then |Option(binomial) \- True Value| ≅ O(ΔR^2 \+ Δt) So you wonder, why should I use a trinomial tree, which is simpler, gives me the same order of convergence? The answer is that there is no reason, other than path-dependent options. If you are path dependent, then it’s sometimes better, because you get 3^n instead of 2^n, more paths to evaluate. ### Exercise How many total nodes are in a trinomial tree with n steps? ## Dividends Most stocks give you discrete dividends. This quarter gave us profit, that we will distribute to our shareholders. We get 20 cents from a $300 share of stock. Typically what happens is that it’s discrete cash, it gets automatically in your account, of this particular date of this particular year, every shareholder. Mathematically, it’s very difficult. We will explain how to deal with cash dividends. The simplest dividend is continuously paid dividends. ### SPY Sidebar (SPY is an ETF, which is done on multiple stocks, which each pay dividends. There are two ETFs, State Street and Vanguard, both of which track the same index. Because they own the stock, they will receive dividends. If they kept the dividends, they would have to pay taxes, which is very complicated. So what they do is distribute the dividends to the ETF shareholders. They calculate. Remember, the weights depend. The way ETF works, is that I own a basket of stocks. The basket is based on (there’s different types, but) in this case it’s *market cap*. You look to the top 500 assets which are in S\&P 500\. Then you look at the largest, say AAPL. The largest component accounts for say 2% of the entirety of my stocks. Market capitalization value for AAPL is $2.5B. The total value of my assets is $50B. If you divide, it gives you 2%. That is my proportion of shares that I will hold in my basket. The reason why they do this is because it’s very easy to calculate the return of the ETF by summing the returns of the components. Now what happens is, obviously the cap value of the company changes, every day, from $200 to $201. If you keep everything the same, the market capitalization value for AAPL grew. Proportionally then, the S\&P 500 portfolio should hold more of AAPL. Obviously that’s not possible (maybe possible with new methodology), but not possible with old. Every day you have to change shares in the big basket. They do this in the middle of the quarter. When the imbecile decided to buy Twitter and call it X. THAT DUDE IS FROM SOUTH AFRICA, A WHITE DUDE FROM SOUTH AFRICA, WHAT CAN HE BE? THAT’S KIND OF RIDICULOUS, RIGHT? When that guy bought Twitter, he bought every share available in the market. The company is automatically withdrawn, it’s a private company one. It was listed on the S\&P 500 and put another company, and then recalculate all of the shares. The whole thing is really simple. Every instrument on the equity market is trivial, they make it sound complicated. What’s a “basis point”? It’s a goddamn percentage.) They take all the dividends, and scale them in such a way that it’s distributed to the owners of the SPY. It’s distributed when they do the rebalancing. When you open an account with any company (Fidelity, Ameritrade, eTrade) you give them your money, but they don’t put your money in your account. THey put it in some fund, which is their own managed fund. And actually, if you open an account, Fidelity ETF that they keep your money in, returns more money than your bank account. So you should buy Fidelity ETF instead of a bank account. Only difference is that if Fidelity goes under, you’re out of luck. Bank accounts get $250K FDIC coverage. ### Continuing… The value of the stock depreciates. $$ dS\_t \= (r \- \\delta)S\_t dt \+ \\sigma S\_t dW\_t $$ The only difference between this and normal is that we have δ which is a known quantity. We do the same basic constructions with r replaced by r \- δ. This is not a big deal. Your function will take an input r, just replace that with r \- δ. ### Known proportional dividend Let’s say $$\\hat{\\delta}$$ is a proportion. ### Elon Rant Let’s say TSLA, because the owner is an imbecile. I never liked this guy. I played games since I got my first computer when I was 12\. They made me think. One of my favorite games is called Path of Exile, which I played since it started 15 years ago. This idiot put out a video where he claims that he’s rank \#5 in the hardcore version of the game. You don’t understand, I played this game for three years before I understood what was going on. It’s really complicated, it’s 10 years old and every 3 months the devs add something new. There’s so much stuff you can do, in ten different ways. Not only that, but you have to literally play it nonstop to be in the top rank. The company is called Grinding Gear game, they created this game after Diablo. Blizzard, the company that issued it, refused to support it, so these three people in NZ made a Diablo-like game. This imbecile puts out a video showing how he managed to be \#5 on the hardcore ladder. Now hardcore means that if you die in the game one time, your character gets deleted. It’s horrible, basically, because if you make one mistake you die, and I’m going to break my computer. This guy basically opens the game and has a tab that says Elon’s Maps. This \*\*\*\* person has obviously paid another dude to play for him, from India or Romania or whatever, and that guy slaved for this guy, for money, clearly, for three weeks nonstop to become the top. And this guy just claims to be that. Why? What’s the point? And obviously people like me who have played for the longest time know that he has no idea what he’s doing. It’s pretty obvious, so what’s the point of doing this. You want to pretend you’re the smartest. I also like chess, I’ve been playing it for a long time. This dude said he was top two in his chess club. He always says top two, because if he said top, someone would say I remember the top. ### Continuing The company value is the total sum of the assets. When you are becoming a public company, you are selling the shares, and if you add the shares together, that’s the total number of assets. (You may also have debt, a liability), but otherwise that’s the value of your company. When you pay someone money (dividends), the value of your company decreases This dividend is paid at time τ. The value immediately becomes $$S\_\\tau \- \\hat{\\delta}S\_\\tau$$ Now what happens here, and why it’s a special case, is because this is $$S\_\\tau (1 \- \\hat{\\delta})$$ And if you remember the multiplicative tree, it goes from [Embedded diagram omitted from the text export.] So you can just price this directly. Then you look at the dividend payment at time τ between 0 and T, and construct the tree. Then you drop at time τ by this value, S to $$S(1-\\hat{\\delta})$$, and then keep computing the tree. Now you have a tree that’s kind of broken, like someone hit it with a bat. [Embedded diagram omitted from the text export.] What about American options? In order to do this, you do the same thing as for the other tree, but you look what’s the value coming from the tree after τ, and compare it with before τ. If you have an option, you are not paid the dividend. There are going to be a bunch of places where you have to exercise. What’s happening here? I forgot to mention this, and this is a mistake in the book, too. Look at the value $$(p\_u C\_u \+ p\_d C\_d) e^{-r\\Delta t}$$ When you look back at the [Combinatorial Formula](#combinatorial-formula), you need to multiply by the discount factor. But when you exercise, this is the expected value of the option in the future. When you compare here, you’re going to take the maximum of this value, and as if you had exercised right before τ, which is $$(S\_\\tau \- K)\_+$$ If the dividend is sufficiently large, the optimal time to exercise is right before the dividend payment. Except that, this is a fake example, there isn’t such a thing. ### Discretely paid cash dividends The most complicated, the most common case. We’ll say D is a quantity of cash that is constant regardless of the price of the stock. We’ll assume it’s known. You can’t do anything if the dividend is not known. If you use the same idea, and you look at some time t\_i, and t\_{i+1}, and τ is in between. The problem with this is that I have an S\_u and an S\_d. Let’s say I pay the dividend. The value of the stock in between these two is going to drop by D units. If we do like we did before, and drop S\_u by D and S\_d by D, and look at the next node, we have the lost the recombination property. [Embedded diagram omitted from the text export.] We have this tree that recombines all nice until time τ, then it becomes unmanageable. That’s the problem here. How do we fix this? The trick is kinda interesting. It’s a pretty complicated construction which is in the book. The idea is that we’re going to create a new tree. This new tree will follow $$\\tilde{S\_t} \= \\begin{cases} S\_t & t \> \\tau \\\\ S\_t \- De^{r(\\tau \- t)} & t \\leq \\tau \\end{cases}$$ Normally, you drop every time τ by the dividend. Now I drop every node below before the dividend, discounted by the time back. What’s happening here is you make all these trees before τ, nonre-combining. Then starting from that point, every point is recombining. This is better because it’s easier to handle exponential Big O at the beginning, where the numbers are smaller. The second thing that’s really smart about this is that if you have observed real markets, this is how real markets operate. Let’s say I own AAPL stock, and AAPl says that stock is $100 today, and then they say due to our quarterly earnings we will have 10¢ dividend per share on March 15, then the price will go down by exactly that amount. This is the present value of the dividend. But in the real world, there is no real price, there is a bid and ask spread. If you’re participating in HFT, every single winning team was running a market making strategy, from providing liquidity to the market. If you’re doing a long-dated option with dividends, you can actually treat it as continuously paying. Everything that we learn in class is not IRL. ## Tree for Deterministic Time Varying Volatility The stock price doesn’t follow geometric Brownian motion as Black-Scholes assumes. This is well-known. We take R\_t, under the observed price. $$ dR\_t \= \\mu \- \\frac{\\sigma^2}{2}) dt \+ \\sigma dW\_t$$ We can take some daily time intervals, t\_1, t\_2, … t\_n. Then we take $$\\log S\_t \- \\log S\_{t-\\Delta t}$$ Which is the literal continuously compounded return, which should be equal to this quantity in the formula. That quantity is very simple, because the stochastic part is normal, then you add a constant. Then your returns, say $$r\_i \= \\log S\_i / s\_{i-1}$$ r\_i… r\_n These should be normal, you can test with a QQ plot. But it’s not normal, it is leptokurtic, it has fat tails. You can try a different variability. The simplest thing to do is σ√t. You will look at two different data points. It will still be normal, but we will make the variance $$\\sigma^2(t \+ \\Delta t) \- \\sigma^2(t)$$ What’s happening here is that some of the observations are coming from one variability, and others are coming from another. If this is possible, then it’s possible to be leptokurtic, combining N(0, 1\) and N(0, 2). That is one way in which you can make this leptokurtic distribution appear. Having said that, the thing about this is the resulting process is not truly random. What happens is, and there’s a formula of this, if you price something from 0 to T, and you have the σ which is a function of time, it is exactly the same as if you are pricing under GBM with a variability equal to the average (integral) of the function, this proof is in Shreve. ## My Research **Quadrinomial tree**. Tree with four nodes. Somehow, somebody has estimated σ(t) and r(t), risk-free interest rate changes in time, and standard deviations changes in time. Take Δt \= T/n. At times iΔt, we take the volatility values $$\\sigma\_i \= σ(iΔt)$$ and $$r\_i \= r(iΔt)$$ and we will also define $$\\mu\_i \= r\_i \- \\frac{\\sigma\_i^2}{2}$$ We take fixed ΔX\_u and ΔX\_d. We will do an additive tree. They are fixed for every t so that the tree is recombining. When I solve my tree, initially, for binomial and trinomial, on the right hand side, we calculate the value coming from the continuous process, which was the r \- sigma squared thing, same for every Δt. Now it’s different because the r and σ are changing, and at that particular time I have to match a different value. So now the intervals have to be changing, because I don’t want my X to change. [Embedded diagram omitted from the text export.] You can see that the probabilities are different. And actually at each step we will assume that the probabilities are the same within the stpe. If you don’t ,then that’s stochastic volatility. For convenience, we will define $$p\_u^i \= p^i, p\_d^i \= 1 \- p^i$$ I also don’t want to deal with ΔX like this, so it will be the same ΔX and \-Δx. So now we have our equations $$ p\_i \\Delta x \- (1 \- p\_i) \\Delta X $$ On the right hand side, it’s supposed to be the integral of an interval, and technically you can pick any point, but the left point will give us the stochastic process. $$p\_i \\Delta X \- (1 \- p\_i) \\Delta X \= μ\_i \\Delta t$$ $$p\_i \\Delta X^2 \+ (1 \- p\_i) \\Delta X^2 \= \\sigma\_i^2 \\Delta t \+ \\mu^2\_i \\Delta t^2$$ The system has a ton of unknowns, because it’s for every i. So there are actually 2n such unknowns. How do we solve it? I’m going to do a thing. We’re going to express $$p\_i \= \\frac{1}{2} \+ \\frac{\\mu\_i \\Delta t}{2 \\Delta X}$$ You can extract the ΔX directly to get $$\\Delta X^2 \= \\sigma\_i^2 \\Delta t \+ \\mu\_i^2 \\Delta t^2$$ We have a problem. You can see here that the μ here changes at every i, but ΔX can’t change. What you do, is you change this, this model doesn’t work. The next step is to make every step change, and then take an average to get ΔX $$\\Delta t\_i$$ Then $$\\bar{\\Delta t} \= T/n$$ and $$\\Delta X \= \\sqrt{\\bar{\\sigma^2} \\bar{\\Delta t} \+ \\bar{mu^2} \\bar{\\Delta t}^2}$$ The point of these bars is the average. $$\\bar{\\sigma^2} \= \\frac{1}{n} \\sum\_{i=1}^n \\sigma\_i^2$$ etc. You take these averages and put them here. The constructions can get complicated. But if you take the binomial tree with these values, you get a pretty good approximation for the moving volatility and the moving interest rate. It’s given to you that σ(t) \= 0.4t \+ sin(t \- π), or whatever. It’s deterministic. ## The most complicated part The problem is that none of these trees solve the Heston model (e.g.) Heston: $$\\frac{dS\_t}{S\_t} \= r dt \+ \\sqrt{y\_t} dW\_t$$ where $$dy\_t \= \\alpha(\\bar{y}- y\_t) dt \+ \\sigma \\sqrt{y\_t} dZ\_t$$ First of all, the construction I have applies to these types of stochastic processes. There are two conditions in the original paper. It can be any general function, for this construction to work, by the way. The stock price must look like this, where it’s dS\_t/S\_t, so when you apply logarithm it becomes explicit, just y\_t. The Brownian motions must also be independent. I had one student that worked on this, and I’m pretty sure it’s solvable, but he quit to go work at Millennium. The question is, how do you construct a tree for this process? It’s similar. The tree has to be recombining, and match the stochastic process. How it goes, the spread… Remember the condition Δt \> 3σ√t. It’s commensurate with σ. The spread is captured by √y\_t. The problem is the variability, like we saw in the deterministic case, it changes. But unlike the deterministic case where we know, there’s 52 weeks in a year, we can plug 1/52 into this bar bar expression. But the problem is that Z\_t is random, so we can’t get the complete thing. We can only get the distribution, and we can’t do the tree because we can’t match it. So the idea that I had, was the following. I’m going to forget about dY\_t, and work with the distribution of y\_t, and at any moment in time t. I’m going to work at the distribution now. I can observe price, but I can’t observe volatility. Implied vol is a surface. You can get these values which will give you some stupid wrong distribution, but you can think of those values as being a histogram. Then I can simplify it and say it takes values y\_1, y\_2, and y\_3. [Embedded diagram omitted from the text export.] The idea is to have a two-dimensional tree, which R on one axis and y on another. From tis one point, I can form multiple different trees. Each point is a σ. The R is a vertex, and I’m moving to four different points. I could move to a different four points for each σ. For each one of these sets of four points, I can move to a different four on a different plane. It’s kind of like a huge pyramid. There’s a two dimensional pyramid, where you have values with a lot of facets. Obviously this is impossible to calculate because it’s non-recombining. Instead I will take a slice of σ values and do a Monte Carlo, some σ\_1 at t\_1, σ\_2 at t\_2 kinda thing. At each sample, I take a different tree. It’s kinda like a combination of Monte Carlo and binomial tree. I spent a year figuring out that binomial tree works, then sent it to a university and realized that there was a mistake and got 0 \= 0\. Then in 2015, they put it online. The first paper I published is wrong. This is very efficient for pricing. Section 6.11 in the book if you’re interested. ## Bond Options, Caps, Floors, and Swaptions - URL: https://sharifhsn.dev/blog/advanced-derivatives-week-03/ - Structured data: https://sharifhsn.dev/api/posts/advanced-derivatives-week-03.json - Description: We’re going to look at bond options, caps and floors, and everything about swaptions. Depending on time, we may go forward. - Date: 2025-02-13 - Exact published timestamp: 2025-02-13 - Topics: Fixed Income, Interest Rate Derivatives, Bond Options, Black Model, Caps and Floors, Swaptions - Categories: Fixed Income - Source: Advanced Derivatives - Source URL: None ## Interest Rate Derivatives We’re going to look at bond options, caps and floors, and everything about swaptions. Depending on time, we may go forward. Basically these are instruments whose payoff is dependent on the level of the interest rates. Generally, interest rates are more difficult to value than equity, or fixed derivatives. One reason is that the behavior of one individual interest rate is more complicated than that of a stock price. For example, there is some mean reversion behavior. Another reason is that the valuation of many products needs to develop that describes the behavior of the entire zero coupon yield curve. Another reason is that the volatility is different at different maturities on this curve. Typically, it’s not a good assumption that volatility is the same for all maturities. Another characteristic is that interest rates are used for discounting the payoff and also for defining the payoff. ## Bond Options First of all, we’re going to discuss bond options. A bond option is an option to buy or sell a particular bond by a particular date for a particular price. Typically, these bond options are traded over the counter, and they are typically embedded in bonds when they’re issued, to make them more attractive. For example, callable bonds contain provisions that allow the issuing firm to buy back the bond at a pre-determined price at a certain time in the future. Usually they cannot be called for a few years, and the value of the call option is reflected in the yields on the bond. Another example is the puttable bond, which contains provisions that allow the holder to demand early redemption at a pre-determined price in the future. Obviously, such embedded put option will increase the value of the bond, these bonds tend to have a lower yield than a bond without such embedded option. Other examples of instruments with embedded options: loans and deposits. The 5-year fixed-rate deposit issued by some institutions (banks) can be redeemed without penalty at any time, this is an American put option. You can see this deposit instrument as a bond. Another example might be a bank loan with 5% per annum quote good for the next 2 months. By having this, you have the right to exercise over the next 2 months. Basically there are many OTC bond options, and some embedded options are European. For this particular lecture, we will assume that the forward bond price has a constant volatility \\(\\sigma_B\\). This will allow us to use Black’s model. There are two options for pricing interest rate options. One is to use a variant of Black’s model. Another is to use a no-arbitrage yield curve based model. Here we’re going to look at pricing such option using Black’s model. Such model will assume that the value of an interest rate, a bond price, or some other variable at time T in the future will have a log normal distribution. ## Black’s Model Revisited We will revisit Black’s model. Consider a European call on an asset with strike K and maturity T. Define P(t, T) as the price at time t of a zero-coupon bond expiring at time T paying $1. Sidebar: A martingale is a zero-drift stochastic process. In general, we can say that a variable θ follows a martingale if dθ \= σdW where dW is the Wiener process. In this case, σ may be stochastic. Then you have the martingale property defined as $$ \\mathbb{E}[\\theta_T] = \\theta_0 $$ We can obtain an equivalent martingale result. Assume that f and g are prices of traded securities. We will assume that these prices are dependent on a single source of uncertainty. Furthermore, let’s define the ratio φ \= f/g. In this case, we’re going to call g the **numeraire**. φ will be the relative price of f with respect to g, the numeraire. We are expressing f in terms of units of g. That’s why the security price of g is the numeraire. Let’s assume also that we have volatilities for f and g. Assume the volatilities \\(\\sigma_f\\) and \\(\\sigma_g\\) in a world where the market price of risk is \\(\\sigma_g\\). We can see that the market price of risk is the volatility of g, then the ratio f/g is a martingale for all security prices. Here we’re going to try to review this result. We can consider the process f, and take df as $$ df = (r + \\sigma_g \\sigma_f) dt + \\sigma_f f dW $$ In general, the process followed by the derivative f can be rewritten as df \= μ f dt \+ σ f dW, aka GBM. The value of μ depends on risk preferences. If the market price of risk is 0, then you have df \= r f dt \+ σ f dW. By taking μ \= r \+ λσ, where λ is the market price of risk, or \\(\\sigma_g\\). Then you will end up with the same relationship as before. $$ df = (r + \\sigma_g \\sigma_f) f dt + \\sigma_f f dW $$ $$ dg = (r + \\sigma_g^2) g dt + \\sigma_g g dW $$ Basically this derivation describes the process f/g. The market price of risk will make this process a martingale. We can apply Itô’s lemma here (not showing work…) for ln f. $$ d\\ln f = (r + \\sigma_g \\sigma_f) - \\tfrac{\\sigma_f^2}{2}) dt + \\sigma_f dW $$ $$ d\\ln g = (r + \\tfrac{\\sigma_g^2}{2})dt + \\sigma_g dW $$ If we have these two processes, we can get the differences as $$ d(\\ln f - \\ln g) = (\\sigma_g \\sigma_f - \\tfrac{\\sigma_f^2}{2} - \\tfrac{\\sigma_g^2}{2}) dt + (\\sigma_f - \\sigma_g) dW $$ Therefore, if you combine these terms, and get $$ d(\\ln\\tfrac{f}{g}) = -\\frac{(\\sigma_f - \\sigma_g)^2}{2} dt + (\\sigma_f - \\sigma_g) dW $$ Here again we can use Itô’s lemma to get the process for f/g. $$ d(\\tfrac{f}{g}) = (\\sigma_f - \\sigma_g) \\frac{f}{g} dW $$ Based on our previous definition/specification, f/g is a martingale. Therefore it follows that $$ \\frac{f_0}{g_0} = \\mathbb{E}_g\\left[\\frac{f_T}{g_t}\\right] $$ and $$ f_0 = g_0 \\mathbb{E}_g\\left[\\frac{f_T}{g_T}\\right] $$ Now we can go back to the zero coupon bond price. Let E\_T \= expectation in a world that is forward risk neutral with respect T(t, T) What can we say about some of these values? g is the numeraire, so g\_T \= p(T, T) \= 1, The price of a ZCB at maturity. g\_0 \= p(0, T) Then $$ f_0 = p(0, T) \\mathbb{E}_T[f_T] $$ Furthermore, for a European call option, with strike K maturity T, the price of such option will be given by c. $$ c = p(0, T)\\mathbb{E}_T[(S_T - K)_+] $$ where S\_T is the asset price at time T. Furthermore, let’s define F\_0 and F\_T as the forward price of an asset at times 0 and T. Therefore the bond option price is $$ c = p(0, T) [F_B N(d_1) - KN(d_2)] $$ $$ p = p(0, T) [KN(-d_2) - F_B N(-d_1)] $$ where $$ d_1 = \\frac{\\ln(\\tfrac{F_B}{K}) + \\sigma_B^2 \\tfrac{T}{2}}{\\sigma_B\\sqrt{T}} $$ $$ d_2 = d_1 - \\sigma_B\\sqrt{T} $$ The characteristic of this pricing model is that the bond price and the strike price should be the cash prices, not the quoted prices. Let’s look at an example. ## Bond Options Example (Black’s) Let’s consider a 10-month European call option on a 9.75 years bond with a face value of $1000. What we’re saying is when the option matures, the bond will still have 8 years and 11 months. The current bond price is $960, the strike is $1000, the 10-month risk-free interest rate is 10% per annum, the volatility of F\_B for T=10 months is 9% per annum. The bond pays a coupon of 10% per year, semiannual payments. Therefore the coupon payments is $50, with the payments expected at 3 months and 9 months. Here the accrued interest at $25. Let’s suppose the risk-free interest rates are different for 3 months and 9 months, 9% and 9.5% per annum. This is a direct application of Black’s model. In order to solve this, we can just apply Black’s model. The value we need to calculate is the forward bond price F\_B, the expected value of the bond at some time in the future. At some time in the future, you might have some coupons that are no longer considered in the calculations. Therefore the forward bond price is $$ F_B = \\frac{B_0 - I}{p(0, T)} $$ where B\_0 is bond price at time 0\. I is present value of coupons in (0, T) In the next 10 months, we are expected to have two coupons. One in 3 months, and one in 9 months. The current cash price is given B\_0 \= 960 Then we can calculate the forward $$ F_B = \\frac{B_0 - I}{p(0, t)} $$ small t in our case. $$ I = 50 e^{-0.25 \\times 0.09} + t0 e^{-0.75 \\times 0.095} = 95.45 $$ We are using the time of discounting (3 months and 9 months), the coupon payment (50), and the interest rate for each coupon (0.09 and 0.095). Forward bond price therefore is $$ F_B = (960 - 95.45)e^{0.10 \\times \\tfrac{10}{12}} = 939.68 $$ Since the 10 month interest rate is 10% Let’s say it’s under specified. The strike price is 1000, but this problem does not specify if the strike price is the cash price to be paid for the bond or the quoted price. For the payment, if you have the quoted price, the value of the bond is going to be quoted price \+ accrual (dirty price). We can investigate these two cases. 1) If the strike price is the cash price that would be paid for the bond on exercise. F\_B \= 939.68; K \= 1000; \\(p(0, T) = e^{0.10 \\times \\tfrac{10}{12}} = 0.92\\); \\(\\sigma=0.2\\) By applying the formula for the Black’s model for the call (not writing it all down) Should be $9.49. 2) If the strike price is the quoted price. We have one month accrual interest that must be added to K, because the most recent payment was at 9 months, and the expiration is at 10 months. Therefore K \= 1000 \+ 100 \* 1/12 \= 1008.33 The rest of the characteristics are the same. c \= $7.97 ## Standard Deviation of ln B Stdev will rise after 0, then fall. ## Forward Bond and Forward Yield The volatilities quoted for the bond options are many times the yield volatilities and the price volatilities. We can use the duration concept. Let’s suppose if D is the modified duration of the forward bond price at option maturity, then the relationship in the change of the forward price and the change in forward yield is given by $$ \\frac{\\Delta F_B}{F_B} \\approx -D \\Delta y_F $$ ## Yield Vols vs Price Vols (Equation 28.4, page 652\) ## Caps and Floors These are interest rate caps and floors. Basically, these instruments, you can think of them as insurance, protection against the increase or decrease in interest rate. They’re going to cap such an interest rate. First of all, let’s consider a floating rate note where the interest rate is reset periodically to a floating rate. We have such reset periods. The time between the resets is known as the **tenor**. Let’s look a little bit at the mechanics. Assume that the tenor is equal to 3 months. - The interest rate on the note for the first 3 months is equal to the initial rate for 3 months, as observed at t \= 0 - The interest rate on the note for the next 3 months is set equal to the 3-month floating rate prevailing in the market. The interest rate **cap** is designed to provide insurance against interest rate on the floating rat rising above a certain level (“cap rate”). For example, assume a principal of $10 million, tenor is 3 months, life of cap is 5 years, cap rate is 4%. This is an important characteristic for the caps, because the payments are made quarterly, so the cap rate is expressed with quarterly compounding. Furthermore, let’s assume on a particular reset date, we have the 3-month floating rate as 5%. In this case, a floating rate note would require a payment of the period (0.25) times the rate (0.05) times the principal (10M) \= $125K. However, with the 3-month rate capped at 4%, it resolves to $100K. The cap provides protection, a payoff of such difference. If you have such cap, then the cap provides a payoff of $25,000, 3 months later So what is the value of the cap? At each reset date that is happening during the life of the cap, the floating rate is observed. If the floating rate is ≤4%, which is smaller than the cap rate, then there is no payoff from the cap. Furthermore, if it’s \>4%, then the payoff is the excess rate applied to the principal. In this case, we have 19 reset dates (5 years \* 4 tenors per year \- 1 initial tenor), at times 0.25, 0.5, …, 4.75 years. There are 19 payoffs, which occur a tenor after, so at times 0.5, 0.75, …, 5 years. ## Portfolio of Interest Rate Options **The cap can be expressed as a portfolio of call options** **on a reference floating rate** (LIBOR maybe?). Consider a cap with a total life T, principal L, and cap rate R\_K. The reset dates are t\_1, t\_2, \\ldots t\_n and t\_{n+1} \= T. We will define R\_k as the floating interest rate for a period between t\_k and t\_{k+1} observed at time t\_k. The small k can take values 1 ≤ k ≤ n. The cap leads to a payoff at time t\_{k+1}, based on the tenor. The payoff depends on the realization of the floating interest rate. $$ \\delta_k L(R_k - R_K)_+ $$ This is applied to the principal. It’s annual, quarterly compounding, so this particular payoff will be applied for the period corresponding to the tenor, which we can mark with the delta. $$ \\delta_k = t_{k+1} - t_k $$ You may assume that tenors are equal, even if that’s not necessarily the case. We can also consider it as a portfolio of puttable bond options, where the ZCB has payoff occurring at time t\_k. The previous payoff at time t\_{k+1} is equivalent to $$ \\frac{\\delta_k L}{1 + R_k S_k} (R_k - R_K)_+ $$ We can use algebra to determine that this is equivalent to $$ \\left(L - \\frac{L(1 + R_K \\delta_k)}{1+R_k \\delta_k}\\right)_+ $$ Therefore the denominator is equivalent to the value at time t\_k of a ZCB that pays L(1+R\_k δ\_k) at time t\_{k+1}. The floor is just the reverse of this, using put options, and R\_K \- R\_k. We can then view each call option in the portfolio that represents a cap as a **caplet**. We assume that the interest rate underlying each caplet is lognormal. Then we can use Black’s model. A collar is an instrument that guarantees that the interest rate on the underlying floating rate always lies between two levels. You can achieve this by taking a long position in a cap and a short position in a floor. In terms of valuation… ## Collar Example The value of a caplet is $$ L\\delta_k p(0, t_{k+1}) [F_k N(d_1) - R_k N(d_2)] $$ using the standard d1 and d2 for option valuation. Where F\_k is the forward interest rate for a period between t\_k and t\_{k+1} In addition, we have to take some assumptions or make some estimates about the volatility of the forward interest rates σ\_k \= volatility of the forward interest rate Here you value all the caplets in order to determine the value of a cap. Consider a contract that caps the LIBOR interest rate on $10M at 8% per annum (quarterly compounding) for 3 months starting in 1 year. This is a caplet, and could be an element of a cap. LIBOR/swap zero curve is flat at 7% per annum, volatility of the rate is 20% per annum continuously compounded zero rate for all maturities is 6.9395% Let’s consider our inputs $$ t_k = 1, t_{k+1} = 1.25, \\delta_k = 0.25, F_k = 0.07, R_k = 0.08, L=10M, \\sigma = 0.2 $$ We need to calculate the discount factor and all the other values $$ p(0, t_{k+1}) = p(0, 1.25) = e^{-0.069395 \\times 1.25} = 0.9169 $$ Furthermore, to calculate for our options $$ d_1 = -0.5677 $$ $$ d_2 = -0.7677 $$ Then we can use the option valuation Black’s model to get the final value ### Assumptions We must use spot volatilities, where the vol is different for each caplet, or one volatility, where it’s flat and the same for each cap. ## Swaptions The swaption or swap option gives the holder the right to enter into an interest rate swap in the future. There are two kinds, the **right to pay** a fixed rate and receive LIBOR, or the **right to receive** fixed rate and pay LIBOR. (LIBOR \= floating) We have a single option on the swap rate with repeated payoffs. This is a series of cash flows $$ \\frac{L}{m}(S_T - S_k)_+ $$ where L \= principal m \= frequency per year S\_K \= swap rate S\_T \= swap rate at time T The value of the swaption on the swap rate with repeated payoffs where the holder has right to pay S\_k. Each of these payoffs will be discounted to the present. $$ \\sum_{i=1}^{mn} \\frac{L}{m} p(0, T_i) [S_0 N(d_1) - S_K N(d_2)] $$ S\_0 is the swap rate at time 0\. This is a natural extension of Black’s model, just with more discount factors. This represents discount factors for mn payoffs ⬇️ $$ \\sum_{i=1}^{mn} p(0, T_i) $$ We can define A as the value of a contract that pays 1/m at times T\_i (1 ≤ i ≤ mn) With this notation, the value of the swaption will become $$ LA [S_0 N(d_1) - S_K N(d_2)] $$ ## Swaption Example Suppose that the LIBOR yield curve is flat at 6% per annum with continuous compounding. Consider a swaption that gives the holder the right to pay 6.2% in a 3-year swap starting in 5 years. The volatility of the forward swap rate is 20%. The payments are semiannually and the principal is $100 million. Let’s set our variables For A, this is the sum of the discounted swaps. The yield curve is flat, so it’s always 6%. We are semiannual, so we start six months after 5 years when the swap starts. $$ A = \\frac{1}{2}(e^{-0.06 \\times 5.5} + e^{-0.06 \\times 6} + \\ldots + e^{-0.06 \\times 8} = 8 $$ We also need S\_0, the forward swap rate. \[...skips some steps, forgot to pay attention) $$ R_c = m \\ln (1 + \\frac{R_m}{m}) $$ Therefore S\_0 \= 0.0609 S\_K \= 0.062 T \= 5 σ \= 0.2 $$ 100M \\cdot 2.0035[0.0609 N(0.1836) - 0.063 N(-0.2636)] = \\$2.07M $$ ## Sensitivity All these derivative have a delta, the DV01, the impact of a 1bps parallel shift in the zero curve. ## Next Lecture We will be doing adjustments, like convexity, time, and quant adjustments Caps protect you, if the volatility rate is too high, you might have risk for a longer period. The swap which is fixed for floating, you can enter in the future. Once you enter the swap, ## Tree Approximation and American Options - URL: https://sharifhsn.dev/blog/computational-methods-week-03/ - Structured data: https://sharifhsn.dev/api/posts/computational-methods-week-03.json - Description: We’re not going to use the arbitrage method, we are going to use the forest methods. - Date: 2025-02-11 - Exact published timestamp: 2025-02-11 - Topics: Computational Methods, Binomial Trees, Option Pricing, American Options, Barrier Options - Categories: Computational Methods - Source: Computational Methods in Quantitative Finance - Source URL: None ## Textbook \- MF3 ## Tree Approximation We’re not going to use the arbitrage method, we are going to use the forest methods. ### Stochastic Process You all understand the stochastic process. Fundamentally, it starts from a point \\(S_0\\), and has these continuous paths that are very weird, because they’re not differentiable. If you blow up the curve, it should smooth out, but actually it stays just as jagged. Therefore we draw an approximation where we draw lines on tiny intervals to approximate the Brownian motion. The thing is, this is a diffusion process. You kind of know the distribution of the path, the \\(dS_t\\) follows \\(\\mu S_t dt + \\sigma S_t dW_t\\). This is an exponential of a normal, it’s called **lognormal**. This distribution has a known shape. A whole bunch of paths will end up at the end. ### Discretization The tree approximation says that the idea that this is continuous is bullshit. So let me do something else. Instead of something that goes all over the place, I’ll start at \\(S_0\\) and approximate the path by discrete intervals. At each step, the distribution should approximate the log-normal thing. It’s the same at each slice. How is this working if the tree is discrete? But as the Δt shrinks, you will have more and more points, you take the interval to be smaller. Each point will have a certain probability, which will reflect the target distribution. ### Kolmogorov In probability theory, there’s a big theorem Kolmogorov. These continuous time processes, as long as you have two, that match in their discrete time distribution. And if the distribution is the same, jointly, you can say it’s the same process. And that’s why this worked. That’s the idea of any tree construction. You want to make it so that it matches each distribution. There are two different requirements. It has to match this distribution, and it has to be **recombining**. Recombining tree will go from one point to two points, the binomial tree. If you don’t recombine, you go from 4 to 8 to 16, n steps will create \\(2^n\\) paths, each distinct. If you do 64 steps, then you use all the bits in your computer. That means your computer will crash. 64 steps is very little, so the alternative is… Recombining! You need to store the nodes, so it will grow at a polynomial rate, \\(1 + 2 + 3 + 4 + \\cdots\\), which is \\(n(n+1)/2 \\to O(n^2)\\), this is much better than \\(2^n\\). The number of paths are the same, you just use the same points. Even though I basically have the number of paths, and I can store them in a computer, I still don’t do anything if I’m pricing path dependent options. There’s a problem: what if this isn’t a Markov process? If it doesn't just depend where I am now, I have to remember where I was before, and that costs 2^n. So this only works when you’re trying to approximate a Markov process, where you can assume that all information is encoded in the value at the end. Sidebar: This is a difference between stochastic processes and LLMs. LLMs are just the most horrible things, they do exponential stuff, because the language is finite, not like numbers. There’s a discussion in the textbook: Florescu’s work is a general diffusion approximation method (6.9) ### The Traditional Tree We’ve done the u and d before. But is that the only tree? It’s just a single tree. I’m not going to do that tree, because that tree is complicated. There are two trees you can make. There is the process $$dS_t = \\mu S_t dt + \\sigma S_t dW_t$$ Everything we’re doing is the geometric Brownian motion, we are not approximating any other stochastic processes. You could, but they think we’re stupid, so they only teach you one single tree of one type. Trees are fascinating. You can do whatever you want, but you need to understand how they work in order to use them. What is the difference between **additive tree** and **multiplicative tree**? $$dS_t = \\mu S_t dt + \\sigma S_t dW_t$$ But then you need to do Girsanov to use risk-neutral measure $$dS_t = r S_t dt + \\sigma S_t dW_t^Q$$ But this stochastic process is multiplicative. If you take an approximation and get $$ S\_{t+\\Delta t} \- S\_t \= \\int\_t^{t \+ \\Delta t} rS\_t dt \+ \\int\_t^{t \+ \\Delta t} \\sigma S\_t dW\_t$$ This is a Riemann integral and stochastic integral. We can approximate this by taking the left point, the value in t. If you do the homework, if you use a quadrature rule, we’re using the rectangle here. Here we’re using the stochastic integral approximation, because we know that the expectation of the increment of Brownian motion is 0\. $$ S\_{t+\\Delta t} \= S\_t(1 + r\\Delta t \+ \\sigma \\Delta W\_t) $$ The resulting tree that you use to approximate \\(S_t\\) is called multiplicative. That’s why we took S and multiplied it by u or by d. That value is supposed to be that thing in parentheses, or something that converges to it. What if I don’t care about this, and I want to approximate in the most efficient way possible? If you take \\(R_t\\) to be the logarithm of \\(S_t\\), and apply Itô to \\(R_t\\), you get $$ dR\_t \= (r- \\tfrac{\\sigma^2}{2}) dt \+ \\sigma dW\_t^Q $$ And then if you discretize, $$ R\_{t+\\Delta t} \= R\_t \+ (r \- \\tfrac{\\sigma^2}{2}) \\Delta t \+ \\sigma \\Delta W\_t $$ And this is direct, not an approximation. The next R\_t is the previous one plus something. What is the difference? In the multiplicative case, you go from some S value with some \\(p_u\\) and \\(p_d\\) to \\(S u\\) and \\(S d\\). But, this is equivalent to \\(\\log S\\) to \\(\\log S + \\log u\\) or \\(\\log S + \\log d\\). These two trees are completely equivalent. This is possible because logarithm is a monotonically increasing function, so it can be a one to one transformation. We’re going to do additive because we’re not imbeciles, much easier than multiplicative. The two conditions: 1. The tree needs to be recombining. I can do that as long as I am dealing with a Markov process. That’s why we learn that GBM is a Markov process in 610, because it’s very valuable. 2. We need that calculated approximating derivative value needs to converge to the value calculated using the continuous process. I don’t really care about the continuous process. I’m trying to approximate the option price/premium. As long as it converges to the true value, then I’m fine. ## Convergence There is a value at time t, a random variable. I quantify the value by looking at the probability I reach every single point, where there’s a probability at each point, a discrete distribution, which should converge to the continuous version. \\(R_t^d\\) is the random variable obtained using a discrete tree. Let \\(R_t\\) be the random variable using the continuous time process. A stochastic process (by the way), there are two things. It depends on the path, and also time. It’s technically a function of two things. For each path, you have times, so you can slice the stochastic process by either the path or the time. What I’m talking about here, it’s at some fixed t, with some random variable coming from the tree and from the stochastic process. We want our tree to converge to the stochastic process. The payoff should $$ \\mathbb{E}[e^{-r(T \- t)}\\varphi(R\_T^d)|\\mathcal{F}\_t] $$ (discount factor doesn’t really matter) This value should converge to $$ \\mathbb{E}[e^{-r(T \- t)}\\varphi(R\_T)|\\mathcal{F}\_t] $$ In the theory of random variables, we have tons of convergences. Any convergence will work. The weaker type of convergence is convergence in distribution. It says that you have a series of random variables that converge in distribution to a target as long as the distribution of the random variables converges to the distribution of the target. With this tree, that’s what we are aiming to do. In most trees, you take Δt to go to 0\. We are going to be doing the additive tree. From \\(R_0\\), we can go only one. Either you’ll go to \\(R_0 + \\Delta R_u\\) or \\(R_0 + \\Delta R_d\\), with some probability \\(p_u\\) and \\(p_d\\). Everything is done in terms of Δt. What about the next step? Technically, I can use a different probability. But we have a very nice GBM, with a normal distribution in Brownian motion. As long as we have the same size of the interval, they have the same distribution, and we can use the same probability. $$ R\_0 \+ 2\\Delta R\_u, R\_0 \+ \\Delta R\_u \+ \\Delta R\_d, R\_0 \+ \\Delta R\_u \+ \\Delta R\_d, R\_0 \+ 2\\Delta R\_d $$ Then you notice that the two middle values can recombine. You don’t need to have any relationship between \\(\\Delta R_u\\) and \\(\\Delta R_d\\), and it doesn’t matter. The other condition is that it has to converge to the continuous distribution. We are working with the additive case, which is very easy. So we know that $$ R\_{t+\\Delta t} \- R\_t \\sim N\\left((r \- \\tfrac{\\sigma^2}{2}) \\Delta t, \\sigma^2 \\Delta t\\right) $$ The tree with this R^d\_t, is an increment with only two values. It’s a very simple Bernoulli distribution, with a probability to \\(\\Delta R_u\\) and \\(\\Delta R_d\\). The normal is characterized by mean and variance. So therefore we will advance a general theorem, which was proved in terms of Central Limit Theorem. There’s a French guy who proved it with binomial distribution. You have a number of trials, and a number of successes, and he looked at average number of successes. As n goes to infinity, the distribution of the average minus p and divided by blah blah converges to the standard normal. And that’s what started the whole Central Limit Theorem, and it was extended to different distributions by Laplace. If you decrease the time interval and you look at the probabilities, they will be binomial probabilities, exactly the same. But there’s a fundamental reason why this is happening. Florescu has proven a more general convergence theorem for trees, discusses infinitesimal generator. I’m going to look at the moments of R\_t^d and set them equal for discrete and continuous, to get our probabilities. How do we solve for these four unknowns? ### Mean $$ \\Delta R\_u p\_u \+ \\Delta R\_d p\_d \= (r \- \\tfrac{\\sigma^2}{2}) \\Delta t $$ ### Variance Pretty ugly, so we won’t equate, we’ll just equate the second moment. We’ll use the property that $$ \\mathbb{E}[R\_t^2] \= \\mathbb{V}[R\_t] \+ (\\mathbb{E}[R\_t])^2 $$ Therefore ### Probabilities $$ p\_u \+ p\_d \= 1 $$ ## What else? DeMoivre proved binomial converges to normal. The normal is characterized completely by the mean and variance, so this is all we need. We don’t need to look for anything else. And actually any number of four parameters solved will give us the tree. There’s an infinite number of trees that solve this. In practice… This was actually initiated and proven by finance guys, who all came up with it from complicated arguments. They solve it for one specific case. If you take ΔR\_u \= \-ΔR\_d \= ΔR, where you go up and down by the same quantity. Now you will have in this system, three equations with three unknowns, which creates a unique solution. This is called the **Cox-Ingersoll-Ross** tree, very famous, the first. You can also take p\_u \= p\_d \= ½. The probabilities are constant, but the values may be different. This is the Trigeorgis tree. For CIR, expressed multiplicatively, you get \\(S_t e^{\\Delta R}\\), and \\(S_t e^{-\\Delta R}\\) And to solve it, there’s no mystery. It’s just a system. \\(\\sigma\\) and \\(r\\) are given to you, by implied volatility and the interest rate. \\(\\Delta t\\) is your choice, typically expressed by \\(T/n\\). Sidebar: Technically, some trees are better than others. The one with symmetric values is faster to converge, but in general they have the same order of approximation. It depends, however, on what you’re approximating. Let’s approximate the European option. This is simple because the payoff only depends on the terminal value. Given this, you have that the price of the premium is $$ \\mathbb{E}[e^{-rT} \\varphi(S\_T)|\\mathcal{F}\_0] $$ There is a simple derivation that makes sense to me, and this is why the tree functions. ~~~text \= \\mathbb{E}\[e^{-r(T-\\Delta t)} \\mathbb{E}\[e^{-r\\Delta t} \\varphi(S\_T)|\\mathcal{F}\_{T-\\Delta t}\]|\\ ~~~ This uses the tower property of conditional expectations. As long as you discount back by something bigger, then you can get it back. Now you can see that this is conditioned on the penultimate step. And you can continue to nest these expectations so that each expectation is conditioned on each step. What this tells you is basically, you can look at the final value \\(\\varphi(S_T)\\) and come back to the previous step, calculate the expectation there, and go on. Then I can calculate the European call if I get \\(S_T\\) from \\(R_T\\). The discount factor is actually irrelevant because either you do it at first, or you do it every step, or at the end, doesn’t matter. You will need to research and construct this algorithm. ## Algorithm inputs: - r: interest rate - σ: volatility - S\_0: initial stock price - T: time to expiration - K: strike price - n: number of steps Then decide what tree you’re going to use, CIR or Trigeorgis $$ R\_0 \= \\log S\_0 $$ $$ \\Delta t \= T / n $$ Then you calculate ΔR. The formula is in the book, based on the tree. Same for p\_u and p\_d. #### What makes us special? I solve problems that no one else can solve, I know math and coding. Nonstandard pricing comes in, it’s the Stevens student that says I can do it. Then you calculate the terminal values at the end. There are n + 1 points after n steps. You have to create a vector of values \\(R[0] = R_0 + n\\Delta R\\), and \\(R[1] = R[0] - 2\\Delta R\\) An algorithm is not something you just wake up and do, you have to think about it on paper, then code it. That’s why we have immigrants. You’re working with vectors, that have n \+ 1 values, if you step back and take the two values and There’s no book with this information. ### Sidebar: Nash Equilibrium If everyone in the room has one objective, and they want to maximize their profit, they all have to interact. The shares are being handed out, you get a mini-market, where they exchange assets to do this. This guy Nash proved that if we all have an objective, there is only one way in which the wealth is distributed that maximizes everyone’s objectives. The theory doesn’t work because it assumes that everyone has the same objective all the time, which is not understanding human beings. If the Nash equilibrium existed, there would be no trading, nobody would desire to have more. ## American Options Other than some horrible equations, you have to use trees. Monte Carlo doesn’t work, because you would like to exercise at the maximum value, but you don’t know that at the time, only when you get to the end. If you are in the tree, you will compare with the expected future value. The discount factor needs to be done in the tree here because you will need to know what the discounted expected value is. You do it from the point and you calculate the value. You have two values, and you take whichever one is higher. This is a very useful way of thinking, related to reinforcement learning. This is based on the idea that you take the best action now which maximizes my future expected reward. Physics is baby shit, with stupid matrix algebra. Finance is way harder. You can also use trees to approximate Heston or SABR. There’s one thing to mention. ## Call vs Put There’s an argument in the book about put call parity. The American call and the European call are the same exact number. This is a mathematical thing because of the stochastic process we do, which is continuous. There was no place in the tree when you stepped back when it was optimal to exercise. But in the put, there are places in which it is appropriate to exercise, which creates a larger number than European put option. Dividends are also a different question, or if the stochastic process jumps, which creates value for the American call. Pretty much every path-dependent option can be done through this trick. The fundamental thing is the expected value of the future, conditioned on the particular node you’re in, which is what you’re storing. HW1 is not about trees, HW2 is about trees. I’m going to ask you to price a barrier option. ### Barrier Option There are two types, the IN and the OUT type. IN option (can be call or put), is worthless unless the stock price hits the barrier. At that point, the option activates. An UP and IN call option, means the barrier is up from the starting price, IN means that once the barrier is reached, then the option is activated, and it’s a call option, which has its own strike price. The OUT option starts as a regular option, which becomes worthless if the barrier is hit. Let’s say you go the Bank of America and say you want to buy options. Instead of having a fixed rate on my house, I want to buy variable rates. Then BofA looks at instruments created by a quant, which gives the rate. That’s an option that’s provided by a specialist. Options are issued by the owner of a stock. You can write covered options if you own the stock. You can write an option if your broker allows it, or if you won the call. You can provide naked calls, where you don’t have the actual share. I have 100 shares up for bidding. I could, using the money I bought from an option, buy another option, and close my contract no matter what. Let’s say I hold onto my option until expiry, and receive my money, and then my share is gone. The broker will settle the contract. The other thing you should look at is, if you look at this data, you should explain the value of the options. There’s something called open interest (openint) and volume. Volume is the number of contracts being transacted, typically per trading day. Open interest is how many contracts are actually open. The number of contracts open is usually very small compared to transacted. The same contract is being bought and sold, 50 issuances traded a thousand times. This is relevant because it applies here. Basically, sometimes you don’t want to have the contract becomes active. There’s no risk of the option cannot be exercised. Barrier options are typically used for fixed income instruments. The barrier technically here depends on the underlying, which for fixed income is the rate, which is floating. For that, you want to have protection because you don’t want to pay more than e.g. 5%, but you don’t want to activate that unless the thing goes crazy. That’s when you put the barrier in place. \\(S_0\\) the contract is either going to be a call, with \\(S_T - K\\) or a put with \\(K - S_T\\) This is not very relevant because the option will become active no matter what. But the barrier matters. DOWN, UP, IN, and OUT are the four parameters, with four possibilities. There is a relationship between this. Let’s say the barrier \\(= B\\). Let’s talk about a CALL UP and IN. This option only becomes valuable when the barrier is hit from below. YOu can model this with an indicator $$\\mathbb{I}\_{\\{S\_t \> B\\}}(S\_T-K)\_+$$ What about UP and OUT? If this is happening, then I get zero, otherwise you get the payoff $$ \\mathbb{I}\_{\\{S\_t \\leq B\\}} (S\_T \- K)\_+ $$ But these two sets are complementary to each other. If you add this using probability theory, you will see that these sum up to 1\. So if you take a call that’s up and in, and a call that’s up and out, you will get a regular call. That means that you don’t have to calculate both of the barrier options, you only need to calculate the regular call and the one that you figure this out. The reason this matters is because the OUT is way easier to price than the IN. The reason is because we are using a tree. If you look at a path in the tree, this is complicated because you have to keep track of the paths. The difficulty is with the point which is under the barrier, but it could be reached by crossing the barrier. If it’s IN, then it becomes valuable. So you have to look at all the paths that cross the barrier, so complicated. But with the OUT, if it passes the barrier, then you set to zero, and don’t count them at all. ## FX Parity and Forward Trading - URL: https://sharifhsn.dev/blog/fx-parity-and-forward-trading/ - Structured data: https://sharifhsn.dev/api/posts/fx-parity-and-forward-trading.json - Description: The FE-635 notes begin with the distinction between spot and forward FX. The forward rate is the exchange rate implied by borrowing in one currency, lending in the other, and carry… - Date: 2025-02-10 - Exact published timestamp: 2025-02-10 - Topics: FX, FX Parity, Forward Contracts - Categories: FX - Source: FE-635 \| Risk Engineering - Source URL: None The FE-635 notes begin with the distinction between spot and forward FX. The forward rate is the exchange rate implied by borrowing in one currency, lending in the other, and carrying the position to maturity. The notation in the class workbook uses the forward FX rate as \(S\), so the quote convention has to be fixed before a formula is evaluated. Covered interest parity is the no-arbitrage check. Domestic and foreign discounting rates determine the forward adjustment, while day-count and currency conventions determine how those rates are applied. A forward is a contract on a future exchange, not a forecast of where spot must end up. The practical checklist is simple: identify the domestic currency, foreign currency, spot quote, maturity, and compounding convention before comparing two prices. ## Yield Curves and Bootstrapping - URL: https://sharifhsn.dev/blog/advanced-derivatives-week-02/ - Structured data: https://sharifhsn.dev/api/posts/advanced-derivatives-week-02.json - Description: We started to look at interest rates. In the first part we’re going to discuss various interpolation methods, functions that are useful for modeling the yield curve. - Date: 2025-02-06 - Exact published timestamp: 2025-02-06 - Topics: Fixed Income, Yield Curves, Interpolation, Bootstrapping, Nelson-Siegel, Smith-Wilson - Categories: Fixed Income - Source: Advanced Derivatives - Source URL: None ## From Last Time We started to look at interest rates. In the first part we’re going to discuss various interpolation methods, functions that are useful for modeling the yield curve. That’s useful because many times you might have to discount or use forward rates or discount factors for maturities that are not available. For this reason you have to do some interpolation. ## Yield Curve Interpolation Methods What do we want to achieve? In general, the interpolation problem can be expressed in the following format: - given some data as a function of time t1, 2… tn and x1, x2.. xn known - construct a continuous function \\(x(t)\\) that satisfies \\(x(t_i) = x_i\\) for \\(i = 1, 2, \\ldots, n\\) What does it mean that the function is continuous? It means that at all points, the limit on left and right is the same, and that it passes through all points. ## Criteria for choosing the Interpolation Method - Smoothness of the forward curve is desirable What does this mean? The derivative should be continuous i.e. it should be differentiable. If the model has smooth curves, it should be more accurate because there’s no reason to have rough changes in curvature. - How local is the interpolation method? If an input changes, does the interpolation function change locally or globally? Simple interpolation, linear interpolation have a local impact. Typically, each input just changes It’s good to have interpolation that impacts globally, but maybe not with very high sensitivity. - Are the forwards not only continuous, but also stable? The **degree of stability** is the maximum basis point change in the forward curve given a basis point change in move in input (???? look this up) ## Linear Methods Let’s graph t against r For \\(t_{i-1} < t < t_i\\) (this is the general case for local interpolation $$ r(t) = \\frac{t - t_{i-1}}{t_i-t_{i-1}} \\cdot r_i + \\frac{t_i-t}{t_i-t_{i-1}} \\cdot r_{i-1} $$ This is a simple rate from drawing a straight line between two points and picking the rate that corresponds to time What is important for this class is to have consistent curves, to prevent arbitrage. If this is the rate, how do we get the forward? The price of a zero coupon bond is the continuously discounted forward rate $$ P(0, t_1, t_2) = e^{-f(0, t_1, t_2) \\cdot (t_2 - t_1)} $$ Then the forward rate is just the natural log of this scaled to time $$ f(0, t_1, t_2) = -\\frac{\\ln P(0, t_1, t_2)}{t_2 - t_1} $$ Zero coupon bonds can also be determined in this way. $$ f(0, t_1, t_2) = -\\frac{\\ln P(0, t_2) - \\ln P(0, t_1)}{t_2 - t_1} $$ In the limit, we can use the derivative $$ f(t) = - \\frac{d}{dt} \\ln P(t) - \\frac{d}{dt} [r(t) \\cdot t] $$ So this is the relationship between the forward and the zero rate r(t) Now we can determine a consistent forward. Therefore f(t) is the derivative rate of this product $$ f(t) = r(t) + tr'(t) $$ So if you take such derivatives, you can substitute the formula for the instantaneous forward. $$ f(t) = \\frac{(2t - t_{i-1})r_i + (t_i - 2t)r_{i-1}}{t_i - t_{i-1}} $$ This is linear, so you can do a function which is linear in log rates if you prefer ## Linear in log rates $$ \\ln(r(t)) = \\frac{t - t_{i-1}}{t_i - t_{i-1}} \\ln (r_i) + \\frac{t - t_{i-1}}{t_i - t_{i-1}} \\cdot \\ln(r_{i-1}) $$ So this is kind of a linear interpolation between two points, but in general these curves are not linear, they have some curvature ## Exponential Interpolation The discount factors are $$ d(t_1) = e^{-r_1 (t_1 - t_0)} $$ $$ d(t_2) = e^{-r_2 (t_2 - t_0)} $$ This is exponential interpolation for the discount curve, which will give us $$ r_1 = -\\frac{\\ln d(t_1)}{t_1 - t_0} $$ $$ r_2 = -\\frac{\\ln d(t_2)}{t_2 - t_0} $$ Then we define $$ r(t_a) = \\lambda r(t_1) + (1- \\lambda) r(t_2) $$ Now we’re going to calculate the discount factor $$ d(t_a) = e^{-r(t_a) (t_a - t_0)} = e^{-\\left(\\lambda r(t_1) + (1-\\lambda) r(t_2)\\right)\\left(t_a-t_0\\right)} $$ $$ = e^{-\\lambda r(t_1)(t_a-t_0)} \\cdot e^{-(1-\\lambda) r(t_2) (t_a - t_0)} $$ ## Cubic Spline Interpolation Given n zero rates of n distinct maturities, we will denote jth maturity and zero rate pair, which we will denote as \\((t_j, R_j)\\) where \\(j = 1, 2, \\ldots, n\\) The idea is to use such polynomials How many can we create? n-1 cubic polynomials. These polynomials will have different coefficients $$ R(0, t) = \\begin{cases} \\beta_{1,0} + \\beta_{1,1} t + \\beta_{1,2} t^2 + \\beta_{1, 3} t^3 & t \\in [t_1, t_2] \\\\ \\beta_{2, 0} + \\beta_{2,1} t + \\beta_{2, 2} t^2 + \\beta_{2, 3}t^3 & t \\in [t_2, t_3] \\\\\\ \\beta_{n-1,0} + \\beta_{n-1, 1} t + \\beta_{n-1, 2} t^2 + \\beta_{n-1, 3}t^3 & t \\in [t_{n-1}, t_n] \\end{cases} $$ That’s the system, but how do we solve it? What might be the problem with this? Right now we have four parameters, and it’s 4 \* (n-1). 4n-4 unknowns, n-1 determinants, so this is an undetermined system. We don’t have information in this system. We can add some more information. The curve is continuous, so we can say that each part should be equal. Skipping some steps, we can say that If R(0, t) is continuous and twice differentiable in t, then n \- 1 equations simplifies to (in a more compact form) observed at time 0, maturity t $$ R(0, t) = a + b(t - t_1) + c(t - t_1)^2 + \\sum_{k=1}^{n-1} d_k (t - t_k)_+^3 $$ This is dependent on the value of t, obviously. Let’s say you have a value of t in between t\_2 and t\_3. $$ \\sum_{k=1}^{n-1} d_k(t-t_k)^3_+ = d_1(t-t_1)^3_+ + d_2(t-t_2)^3_+ + d_3(t - t_3)^3_+ + \\ldots $$ The first two terms will have positive values, and the rest will be 0 Now we’ve restricted our unknowns to a, b, c, and the discount factor \\(d_q, d_2, \\ldots d_{n-1}\\) which is n \+ 2 unknowns with n discrete rates. This is still not fully determined. You may see different results in different implementations based on the next steps, in Python or R. We need two more conditions. It depends on the problem. Typically, there are conditions that make sense. The first derivative at the ends of the intervals, we can make the assumption that the convexity is 0 right at the end of the intervals. $$ \\lim_{t \\rightarrow t_1} R''(0, t) = 0 $$ This counts as two conditions because it’s left and right. Then you have n \+ 2 linear equations and n \+ 2 unknowns. Now we are fully determined. The solution is very easy to obtain 😏 Let’s take the vector R\_1 is going to correspond to t\_1 Then you have to take the second order derivative, so what do you get for c? First derivative is 2c In another way, you could write this as $$ \\begin{bmatrix} R_1 \\\\ R_2 \\\\ \\vdots \\\\ R_n \\\\ 0 \\\\ 0 \\end{bmatrix} = A \\cdot \\begin{bmatrix} a \\ b \\ c \\ d_1 \\ \\vdots \\ d_{n-3} \\ d_{n-2} \\ d_{n-1} \\end{bmatrix} $$ Which we can solve by inverting the matrix $$ A^{-1} \\cdot \\begin{bmatrix} R_1 \\\\ R_2 \\\\ \\vdots \\\\ R_n \\\\ 0 \\\\ 0 \\end{bmatrix} = A^{-1} \\cdot A \\cdot \\begin{bmatrix} a \\ b \\ c \\ d_1 \\ \\vdots \\ d_{n-3} \\ d_{n-2} \\ d_{n-1} \\end{bmatrix} $$ And then $$ \\begin{bmatrix} a \\\\ b \\\\ c \\\\ d_1 \\\\ \\vdots \\\\ d_{n-3} \\\\ d_{n-2} \\\\ d_{n-1} \\end{bmatrix} = A^{-1} \\cdot \\begin{bmatrix} R_1 \\\\ R_2 \\\\ \\vdots \\\\ R_n \\\\ 0 \\\\ 0 \\end{bmatrix} $$ We got the values in A by taking the second order derivative, and the limit approaching t\_1 Explanation of derivatives $$ R(0, t) = a + b(t - t_1) + c(t - t_1)^2 + \\sum_{k=1}^{n-1} d_k(t - t_k)^3_+ $$ $$ R'(0, t) = b + 2c(t - t_1) + \\sum_{k=1}^{n-1} 3d_k(t-t_k)^2_+ $$ $$ R''(0, t) = 2c + \\sum^{n-1}_{k=1} 6d_k (t-t_k)_+ $$ Then the limit will show this collapse $$ \\lim_{t \\rightarrow t_1} R''(0, t) = 2c $$ $$ \\lim_{t \\rightarrow t_n} R''(0, t) = 2c + 6d_1(t_n - t_1) + 6d_2(t_n - t_2) \\ldots $$ But the matrix is a more elegant form. This particular function will pass through all of these points, and it will be a beautiful curve. ## Functional Form of Yield Curve In general, if coupon bonds are available, then use the functional form We’re going to look at three popular models: ### Nelson-Siegel Model This is no longer an interpolation function, it’s more of a fitting function. We will write the specification and the characteristics for this function, and understand the procedure $$ R(0, t) = \\beta_0 + \\beta_1\\left( \\frac{1- e^{-\\lambda t}}{\\lambda t}\\right) + \\beta_2 \\left( \\frac{1 - e^{-\\lambda t}}{\\lambda t} - e^{-\\lambda t}\\right) $$ where we have \\(\\beta_0, \\beta_1, \\beta_2\\) constants and \\(\\lambda > 0\\) to specification. This is finding four parameters. This particular function has some characteristics, we can look at these. $$ \\lim_{t \\rightarrow \\infty} R(0, t) = \\beta_0 $$ If you take a longer maturity, it goes to a constant. So \\(\\beta_0\\) should be associated with the long-term interest rate, like the level of it. Before that, it has some curvature. What about at 0? Well, you would end up a 0/0, so you can apply L’Hopital’s rule. $$ \\lim_{t \\rightarrow 0} = \\beta_0 + \\lim_{t \\rightarrow 0} e^{-\\lambda t} + \\beta_2 \\lim_{t \\rightarrow 0} (e^{-\\lambda t} - 1) $$ $$ = \\beta_0 + \\beta_1 $$ This is the spot interest rate, at t \= 0\. Level Slope Curvature represent the three betas. We have a numerical example just to see #### Example Assume that there are N coupon bonds. For i \= 1, 2, \\ldots N, the ith bond has the following specification B\_i \= cash(dirty) price of bond m\_i \= remaining number of coupon payments These bonds have different maturities, and they may have different specifications in the coupon payments (frequency and amount) t\_i^j \= coupon payment date for j \= 1, \\ldots m\_i t\_i^1 \= date of next coupon payment t\_i^{m\_i} \= maturity date and date of last coupon $$ \\delta_i = t_i^{j+1} - t_i^j $$ P\_i \= principal C\_i (amount of each coupon) \= \\(\\delta_i P_i \\cdot Coupon Rate\\) B\_i \= Market Quote (Clean Price) \+ \\(C_i \\frac{\\delta_i - t_i^1}{\\delta_i}\\) (Accrual of Interest) Theoretical Price: $$ \\hat{B_i} = \\sum_{j=1}^{m_i} e^{-R(0, t_i^j) \\cdot t_i^j} \\cdot C_i + (Then Discount Principal) e^{-R(0, t_i^{m_i}) \\cdot t_i^{m_i}} \\cdot P_i $$ Basically you can calculate this for all bonds. If you want to write this as an optimization problem, we want to find \\(\\beta_0, \\beta_1, \\beta_2, \\lambda\\) that minimizes $$ \\sum_{i=1}^N w_i(\\hat{B_i} - B_i)^2 $$ that satisfies $$ \\begin{cases} \\beta_0 > 0 \\ \\beta_0 + \\beta_1 > 0 \\ \\lambda > 0 \\end{cases} $$ This w\_i item is the weight assigned to bond i, if you want to equally weight them then it’s 1\. Choices for w\_i: - w\_i \= \\(\\frac{1}{t_i^{m_i}}\\) This one is a function of maturity, so longer maturity will have smaller weights. This will fit better on shorter maturity than longer maturity. - w\_i \=\\(-\\frac{1}{B_i} \\cdot \\frac{\\partial B_i}{\\partial YTM}\\) This one is dependent on the duration of the bond ## Siegel-Svensson Adds a second hump term (2 possible maximaminima Let R(0, t) be the zero rate for maturity t, then $$ R(0, t) = \\beta_0 + \\beta_1\\left( \\frac{1- e^{-\\lambda_1 t}}{\\lambda_1 t}\\right) + \\beta_2 \\left( \\frac{1 - e^{-\\lambda_1 t}}{\\lambda_1 t} - e^{-\\lambda_1 t}\\right) + \\beta_3 \\left(\\frac{1 - e^{-\\lambda_2 t}}{\\lambda_2 t} - e^{-\\lambda_2 t}\\right) $$ We have some constraints: $$ \\begin{cases} \\beta_0 > 0 \\ \\beta_0 + \\beta_1 > 0 \\ \\lambda_1, \\lambda_2 > 0 \\end{cases} $$ The optimization problem for the previous problem was for price of the bonds. But otherwise we can write it as We have the factor function, then the fitting function, then the objective function. That’s how you get Nelson-Siegel. For the optimization methods, there are some characteristics, constrOptim.nl is what he uses estimates ## Smith-Wilson Generates a smoothed interpolated and extrapolated term structures that fit the spot market rates. has simplicity, works for low number of market data points but performance This going to work like an interpolation, but it also extrapolates for longer maturities. Parameters: UFR: ultra long-term forward rate LLP: last liquid point, where the zero coupon market support ends (e.g. 20 years) CP: Convergence ## Assignment This homework is due in two weeks. The first problem is similar to what we did last time, given some cash flows from different bonds, fill in the curve. Determine the forward curve, bump it, and then if you change it, what’s the change? If you have 10 years, you have 10 such cases, bump for each year. ## Black–Scholes, Calibration, and Implied Volatility - URL: https://sharifhsn.dev/blog/computational-methods-week-02/ - Structured data: https://sharifhsn.dev/api/posts/computational-methods-week-02.json - Description: Hanlon 2 is occupied by the system administrator, Zheng Xing. He will move his class so that we can occupy there. They will need to remote desktop into the lab. - Date: 2025-02-04 - Exact published timestamp: 2025-02-04 - Topics: Computational Methods, Black-Scholes, Calibration, Implied Volatility, Stochastic Volatility - Categories: Computational Methods - Source: Computational Methods in Quantitative Finance - Source URL: None ## Bad News Hanlon 2 is occupied by the system administrator, Zheng Xing. He will move his class so that we can occupy there. They will need to remote desktop into the lab. I should have come here early to help him 🙁 There’s two modes, Zoom Mode, and Bypass Zoom. If you bypass, you can’t use microphone or camera. Doesn’t really matter, next week will be different. ## Plan We’ll look at some stochastic models. Then we will look at finding roots, and finding volatility, and so on. ## Black-Scholes Model Traditional model: $$ dS_t = \\mu S_t dt + \\sigma S_t dW_t $$ The thing about this equation, is that it is under the objective probability measure. You can see that there is a μ parameter that characterizes the stock. Last time we talked about MLE. That is concerned with estimating the parameters, such as μ and σ. This can only be done when you observe the actual stochastic process prices. It’s an estimation that you do under this objective probability measure. However, when you calculate option prices, you go to the equivalent measure Q which is risk-neutral. The equation is slightly different. $$ dS_t = r S_t dt + \\sigma S_t dW_t^Q $$ We can observe this by doing the Girsanov transformation. So what’s the difference here? The option prices are always under Q. ### Calibration There is a different method to estimate parameters, called **calibration**. We’re going to mention this idea here, and then expand later. *Estimation and calibration are different*. Estimation is a very standard method in statistics. The condition of estimation is that you observe the option prices. Calibration is a term invented in finance, which is only applicable here. It deals with this particular situation where you observe option prices \\(c_1, c_2, \\ldots c_n\\). These are *derivative prices*, not the underlying prices, derived under Q. If you’re estimating parameters for this, you’re going to be estimating r and σ. We’re lucky because under Black-Scholes because σ is the same. But we can’t estimate μ at all, we need stock prices to do that, because μ vanishes. The method is, we would set theoretical prices “equal” to observed prices, and then we solve for parameters. And this is called calibrating parameters. In traditional statistics, you assume that this is a number. But when you do this methodology, and you estimate (e.g.) implied volatility, you would see how it changes based on different option prices. But then you say that σ is different, so how do you adjust for that? When you calibrate to real data, then you get it. The problem is that the model is bullshit. Every time you use derivatives, you are estimating under Q. The reason we are using BS and implied volatility is because there is a formula for this. You should know the formula. C the option price is a function of both S and t, the observed stock price and the time. However each price depends on other stuff, like K the strike price, r the interest rate, and T the time to maturity, and σ the volatility. It looks like there are six things in this equation, but it actually only depends on time to maturity, so there are five things. $$ C(S, T - t, K, r, σ) = SN(d_1) - Ke^{-r(T-t)} N(d_2) $$ Where $$ d_1 = \\frac{\\log (S/k) + (r + \\sigma^2 / 2) (T-t)}{\\sigma \\sqrt{T - t}} $$ $$ d_2 = \\frac{\\log (S/k) + (r - \\sigma^2 / 2) (T-t)}{\\sigma \\sqrt{T - t}} $$ And N is the normal cdf. If you know K and you divide by K everywhere, you can get the same equation in terms of **moneyness**, which is determined by S/k. Moneyness is used in practice, because, for example, if you calculate at the money option for AAPL, where it’s $600 at the money, then the following options are $605, then $610. But for AMD, it’s $10 at the money, then $10.5, then $11. By using moneyness, you scale all of these in the same way. S is observed, K is observed, r is not really observed but you keep it fixed. You just go to the Federal Reserve and pick whatever risk-free interest rate, just keep it consistent, 3M or 3Y, whatever. Time to maturity is a small value. Everything here is expressed in years. **You need to keep this consistent: *everything is in years*.** Unit conversions are a big problem. This will give you yearly volatility. This is something well understood. S\&P 500 has volatility 0.2, IBM 0.6. This is relatively stable. You could express everything in terms of days, but then you would get a scaled daily volatility which is very small. Because of this normality assumption. You can get returns here by doing a transformation and applying Ito. To be clear, this is under Q. $$ R_t = \\log S_t $$ $$ dR_t = (r - \\tfrac{\\sigma^2}{2})dt + \\sigma dW^Q_t $$ You always eliminate S this way when you take the logarithm. Look at the form of this\! If you discretize this process, you get $$ R_{t + \\Delta t} - R_t = \\int_t^{t+\\Delta t} (r - \\tfrac{\\sigma^2}{2}) dt + \\sigma (W_{t + \\Delta t} - W_t) $$ You can see the increment of the Brownian motion, which is normal. This is why the return is very nice, because this is the increment. Then you get the continuously compounded return, which is normal and also independent. This is very nice to work with, and it only works for the Brownian motion. ### Continuously Compounded Return Typically if you calculate return over a year, let’s say you have an asset that changes over time. That has the following return: $$ r_t = \\frac{S_{t + \\Delta t} - S_t}{S_t} $$ This is the simple return, and tells you what happens at the end. But what happened in the middle? When you make the increments smaller, you eventually get $$ R_t = \\log \\left(\\frac{S_{t+\\Delta t}}{S_t}\\right) $$ Every moment in time, I’m going to earn more. This is a mathematical construct and doesn’t exist. But it’s convenient, because it gives us the mean plus variance of the normal. ## Approximating PDEs If V is a European type option, you can only exercise at maturity. V(T) is then only a function of S(T). It only depends on the stock price at maturity. You can have European call or put, as long as the value only depends on the stock price at maturity. We can show that V solves a particular type of education, which is called the parabolic PDE. $$ \\frac{\\partial V}{\\partial t} + \\frac{1}{2}\\sigma^2S^2 \\frac{\\partial^2 V}{\\partial S^2} + rS \\frac{\\partial V}{\\partial S} - rV = 0 $$ This is the PDE, but in order to solve this, you need boundary conditions to solve this. The proper solution is a surface, with an infinite number of solutions. We know that $$ S \\in [0, \\infty] $$ $$ t \\in [0, T] $$ Let’s consider the boundary at T, the time of payoff. $$ V(S, T) = f(S_T) $$ We are not necessarily concerned with the function itself, it could be K \- S, we’re not worried about that. In the theory of PDEs, there are boundary conditions expressed in the function itself, or of derivatives of the function. there are also Dirichlet conditions. We can’t put a condition at t \= 0, because that’s what we’re looking for, the price at t \= 0\. We can put a condition on the boundary at S \= 0, and when S converges to ∞. These are specific, and depend on the option you are pricing. If you are looking at a call option, going to infinity becomes S \- K. which becomes 1 if you derive with respect to S. Going to 0 becomes 0 because We will talk about this later. But these conditions are important\! These boundaries involved are [Dirichlet boundary condition](https://en.wikipedia.org/wiki/Dirichlet_boundary_condition) ## Greeks The Greeks involve the derivatives of the parameters. We have five Greeks. K is fixed so it doesn’t have a derivative. $$ \\frac{\\partial V}{\\partial T} = \\theta $$ When time changes, how does the option change? $$ \\frac{\\partial V}{\\partial r} = \\rho $$ The interest rate changes for time to time, but it’s not that big of a deal. $$ \\frac{\\partial V}{\\partial \\sigma} = Vega $$ By this theory, σ is supposed to be constant. So this doesn’t really exist, it’s kind of a joke. $$ \\frac{\\partial V}{\\partial S} = \\delta $$ $$ \\frac{\\partial^2 V}{\\partial S^2} = \\gamma $$ These are the big kahunas. Delta is the sensitivity to changes in stock price. Gamma measures the concavity or convexity, wherever you are on the curve. These are used in hedging. ## Job Market Tangent Fortunately, finance is very competitive. So it’s very hard to find a job based on your specialty, becoming a quant. JP Morgan has 300 quants in total, throughout its whole company, out of tens of thousands. 200 of them are in England. It’s very hard to become a quant. Those are real quants that do this shit. The good news is that this program doesn't only prepare you to become a quant. It prepares you to become a quant in any company. JP Morgan has 300 real quants, but it has data scientists, quantitative analysts, that do other related things. ## Heston Model I’ve been making fun of this stochastic process BS, what is the alternative? Why do we keep learning about IV if it doesn’t really work? Two reasons. 1\. Most people in finance are imbeciles, so BS is the only thing they can understand. Ten years ago IV wasn’t on Yahoo Finance. 2\. It’s a self-fulfilling prophecy, because everyone uses it. I’m going to buy Trump stock because he’s elected, so it goes up. Now there’s lots of people buying, but nobody selling. But the guy selling will see that the model says that they’re overpaying, so he will sell. The Heston model is not useful for vanilla option prices. If you happen to work as a quant, you won’t work with idiots trying to buy Trump stock. You are going to deal with another quant who works at Barclays who says I have $500 million cash flow, and I want to ensure that the rate I’m getting is 5%, so I’m going to do a swap. Give me the price of a swap. This is a completely nontraditional thing, some instrument that you need to price. In order to understand this non-standard situation, you need to understand the simple model. $$ dS_t = rS_t dt + \\sqrt{V_t} S_t dW^1_t $$ $$ dV_t = K(\\theta - V_t)dt + \\sigma \\sqrt{V_t} dW^2_t $$ These two Brownian motions are correlated by ρ $$ \\rho dt = \\mathbb{E}[dW^1_t dW^2_t] $$ You can write a Heston PDE and solve it with the methodology in the book. Let’s just understand this structure. Note that S\_t is linear, so we can get rid of it and make it explicit in terms of the process V\_t and the Brownian motion. We can do the Ito to the logarithm of S\_t. The Brownian motion part can be negative. You can approximate the increment of V\_t by the dt, because the increment is normally distributed. ### Mean Reversion θ \- V\_t will push towards θ because it reverts to it. The increment will be negative if V\_t is above θ, and positive if V\_t is below θ. And the mean of this process is actually θ, which we can calculate. In practice, K is called *the speed of mean reversion*. If the speed of mean reversion determines how much it goes back and forth across the mean. Keep in mind that this is stochastic and not guaranteed, but this is the average trajectory. ## SABR \- Stochastic α β ρ This is one of the first stochastic volatility models to be created, and was used in the insurance industry. $$ dS_t = \\alpha_t S_t^\\beta dW_t $$ $$ d\\alpha_t + \\nu d_t dZ_t $$ ρ correlation This doesn’t have a solution, but it’s very popular. Assume that the stochastic process follows the SABR model. THen you have an option price, but with no formula. But you do have a formula for the IV of that option price. So what you do is calculate that IV, and then plug it into BS. The reason that it was powerful because interest rate models had used BS already in their code, so they would need to change all their code to use something else. They were all spaghetti code idiots. There used to be a guy at JP Morgan who was 70 years old and kept working. The reason was because he wrote code in COBOL 30 years ago and was running $5 billion instruments valued every day based on this model. Jamie Diamond commanded that they rewrite everything in Python, so he’s gone now. ## Cox-Ingersoll-Ross $$ dV_t = \\alpha (\\bar{V} - V_t) dt + \\sigma \\sqrt{V_t} dW_t $$ I like α for speed of mean reversion and \\(\\bar{V}\\) for my mean reversion. The square root thing is called the **vol of vol** parameter. You can use this for stock prices directly. Before they went back to Vasicek, they used this for interest rates, or anything dealing with them like swaps, caps, floors, etc. The question is, what’s the expected value and squared? We can solve this very neatly by expressing this as an integral. $$ V_t - V_0 = \\int_0^t \\alpha(\\bar{V} - V_s) ds + \\int_0^t \\sigma \\sqrt{V_s} dW_s $$ $$ \\mathbb{E}[V_t] - \\mathbb{E}[V_0] = \\mathbb{E}\\left[ \\int_0^t \\alpha(\\bar{V} - V_t) dt \\right] + \\mathbb{E}\\left[\\int_0^t \\sigma \\sqrt{V_t} dW_t\\right] $$ The expectation of the right side is 0 because it’s a Brownian motion. For the left side, as long as what is under the Riemann integral is finite, the integral order doesn’t matter. So we can flip them, and we want to know the expectation. We will define a deterministic function here based on the time t. $$ \\mathbb{E}[V_t] = v(t) $$ Now we can determine $$ v(t) - v(0) = \\int_0^t \\alpha(\\bar{V} - v(s)) ds $$ This is a Riemann integral that we can directly solve. Then take the derivative. $$ v'(t) = \\alpha(\\bar{V} - v(t)) $$ $$ dv = \\alpha \\bar{v} dt - \\alpha v(t) dt $$ α appears out of nowhere if we derive $$ e^{\\alpha t} $$ so we will actually multiply this to make it appear $$ e^{\\alpha t} dV + \\alpha e^{\\alpha t}dt V = \\alpha e^{\\alpha t} \\bar{V} dt $$ Now we can integrate to give us the solution. \[I can’t see the solution from where I’m sitting, check the textbook for this\] The trick is literally useful when you can get some nice mean here in terms of prices. If you can say that, if the stochastic part was deterministic, then it would be easy to solve. V^2 is left as exercise. Tips: Apply Itô, then use the same trick to get rid of the stochastic part. ## Implied Volatility Implied volatility fits somewhere in between the bid and ask price for options. Mathematicians solve this problem with a formula to calculate option price, such as BS. In that model, σ is the only thing we don’t know. In order to solve this, we will set the option price formula to the average of the bid and ask (mid) and then solve for σ. You can rearrange as the difference between the theoretical price and the mid is 0\. Then you can say you’re finding the root of a function. $$ f(\\sigma) = C(\\sigma) - mid = 0 $$ This value is called the implied volatility. There is an obvious problem with this. This is for K\_1 and t\_1. What if I have a different strike price, or a different time to maturity? Then I’m going to have a different value. This is despite the fact that volatility is calculated on the underlying stock price and therefore should be constant. This is called the **volatility smile**, or really the volatility smirk because the upwards trajectory is only in one direction. If you graph moneyness against implied volatility, you get the smirk. This only goes for one time. For T\_2 larger than T\_1, the smile is less clear, it’s shallower. How do you calculate implied vol? You might want to take an average of these points. Or you could take an average of a certain range. Let’s say you have the data for option prices. Let’s do this $$ \\min \\sum_{i=1}^n (C(S, T_i, K_i, r, \\sigma) - C^i)^2 \\cdot W_i $$ Where you weight each price and you minimize the volatility. This is a fix, because the model is crap. There’s something better. ## Local Volatility This was pioneered by a bunch of math guys at Bloomberg. Dupire, Derman (director of FE at NYU, now retired). They said, let’s fit everything. We’ll look at these values and create the local vol surface. It is a bad model, but it’s what people use. The industry uses a combination of local vol and stochastic vol. The professor’s students who work use this. Local vol is not that complicated. Remember the equation for stock price from BS? $$ dS_t = S_t(r - \\tfrac{\\sigma^2}{2})dt + \\sigma S_t dW_t $$ What these guys said, is that the dt term is crap, we’ll say it’s 0\. The σ is clearly not constant, but I don’t want it to be a stochastic process. Then the natural thing to do is replace σ with σ(t) which is a function of time. Now this is a failure because nothing changes. Now they have tons of data, and this is used for interest rates. Interest rates have two important times, the tenor and the maturity. Tenor is how much time remains. So they need a way to fit based on this information. Let’s do σ(t, S\_t). Where do we get this function from? In the dV PDE, we have σ. So let’s extract σ. ~~~text \\sigma(T, K)^2 = 2\\frac{\\tfrac{\\partial C}{\\partial t} + (r_T - 2T)K\\tfrac{\\partial K}{Z} ~~~ This is pretty much the same question in terms of strike price and time. For each particular option price, I can estimate these derivatives. I’m going to market, r\_t, and Q\_t re known. Then you can construct a vol surface from this, which you can find in Bloomberg. The whole thing is bullshit. I wrote a paper about this. You’re fixing the problem by creating another problem. These derivatives are a big problem. How do you solve for the derivatives? You use finite differences (we will mention later when talking about approximating PDEs). If you have a function f, the derivative df/dx is approximately f(x+Δx) \- f(x). As long as you take Δx goes to 0, this will converge to the derivative. This is called the finite difference, which you can express in multiple ways. This is first-order difference, and then you can use it substitute the derivatives. To get the second derivative, you do the first difference of the first difference. You can see that this is the strike price minus a little bit. But the problem is that it DOES NOT EXIST. In practice, strike prices are very far apart from each other. They justify this for use in interest rate models, in which the interest rate increments are very small. But then with interest rates, the change in time is VERY FAR, so each of them will fail. ## Paper This stuff is not new, in 1998\. Black-Derman-Toi? Model. Florescu heard about it and got mad, and wrote a paper to critique this. *Personal reminder: look at this on Canvas.* ## ChatGPT Gradient descent is “Brute force”, so optimization actually SUCKS, brute force is the best thing. ## Bisection If I have f(x), how do I find x\_0 such that f(x\_0) \= 0\. This is finding the root of a function, which is a famous problem. Bisection is very stupid and very simple, and it only works for our problem. It’s designed on the fact that if you have a root on an interval, then you see that f(a) \* f(b+a/2) \< 0, and if it’s true, then you can bisect this interval and look for the root, repeating like binary search. None of this requires the BS formula, you just need a way to get the option value for a set of parameters. You could use an approximation method like a tree or finite difference. But because you are doing this a lot, you would like to have an analytical solution to be faster. A couple of issues with this: - Interval \[a, b\] needs to contain the root. There cannot be two roots. The option price is increasing in terms of σ, and this is monotonically increasing, so there is only one root. The root is always positive, we know that IV is always positive, so you can always use a \= 0\. Idiots tell you IV is always less than 1 because it’s a percentage. But the IV depends on the actual Call option price. If the price is way weird, then it’s not. The bid might not have been traded in a while, and the ask price moved with the stock price, so the spread is now huge. So now the average is way off. If you consistently get the root to be a or b for certain options, it is usually a problem with your data. IV is usually not huge, although it is possible when stock price moves a lot but option prices don’t. If you use \[0, 1\], you won’t capture that. The fix is simple, to use \[0, 10\], or \[0, 4\]. In two steps, you’re at \[0, 1\]. - This only works for R → R. You need to map into R to do comparisons of greater or less than, like Euclidean distance. ### Newton Method $$ x_1 - x_0 = \\frac{f(x_1) - f(x_0)}{f'(x_0)} $$ My goal is to find x where f(x) \= 0, aka the root. Let’s rearrange $$ x_1 = x_0 - \\frac{f(x_0}{f'(x_0)} $$ where we delete f(x\_1) Then I can keep doing this until the two points are the “same” (within ε) This is the recurrence formula. It works for R\_n generalized, but there is a problem. If you start too far away, you’re going to go in a weird direction. ## Trump I saw a conference from Trump on Tuesday when he brought a CEO from Morocco, and the Sam Altman guy. He thinks that because it creates words, it’s alive. Trump said we need to be leaders in AI. “I am prepared to give you $500 billion” The next day Chinese do it cheaply. Next class, approximate stochastic processes using trees. First homework is due… will be determined tomorrow. We will have 2-3 weeks to do it. ## Old Paper ## Introduction We are concerned with estimating expected future volatility. One approach is with historical data. Another is to use market option prices and derive volatility from option valuation formulas. Doing this derivation requires some assumptions based on which formula you’re using. Black-Scholes requires asset prices to follow geometric Brownian motion and constant volatility. Since volatility is constant, then this disagrees with the real-life volatility smile. In order to account for this, we might think to change volatility based on the stock price. This typically means some kind of stochastic volatility model, which is not easy to solve for. In a special case, if the volatility is deterministic, then we can use the Black-Scholes PDE to solve option prices. DK, D, and R have developed “local volatility” which does the following: “Their methods attempt to fit a cross section of option prices and deduce the future behavior of volatility as anticipated by market participants. Rather than give a formula or a structural form for the volatility function, they search for a binomial or trinomial lattice that achieves an exact cross-sectional fit of reported option prices.” In this way it’s similar to implied volatility, but it attempts to capture the volatility surface. ## Bond Pricing, Duration, and DV01 - URL: https://sharifhsn.dev/blog/advanced-derivatives-week-01/ - Structured data: https://sharifhsn.dev/api/posts/advanced-derivatives-week-01.json - Description: Take-home exam format, 24/48 hours, lots of writing code, during finals week. - Date: 2025-01-30 - Exact published timestamp: 2025-01-30 - Topics: Fixed Income, Bonds, Yield Curves, Bond Pricing, DV01, Duration, Convexity - Categories: Fixed Income - Source: Advanced Derivatives - Source URL: None ## Final Exam Take-home exam format, 24/48 hours, lots of writing code, during finals week. ## Fixed Income Instruments If a company wants money, it has options. It can issue stock, get a bank loan, or issue a bond. What is the cost of borrowing money? What determines this cost? One such characteristic is that it can be determined by the market. If more people want to lend money at a certain maturity, it can lower the interest rate, with money being the commodity. This rate is determined by the central bank, which has influence on short-term borrowing. Some governments can determine the cost of borrowing. When a particular entity wants to borrow money, it depends on rating score. There’s a probability they don’t pay back, that’s risk. Inflation also plays a role in long-term expectations. ### Money and Bond Markets The main factor in the variability is the interest rates. The bonds and rates have an inverse relationship. You can see every fixed income instruments as a stream of known (or rather contingent) cash flows. The interest rate is not a single interest rate for a particular maturity. If you look at Bloomberg, you will have many different rates that are associated with the Treasury, individual corporate entities, and so forth. The money market is traded with banks and corporations. Cash flow delivery can be done up to 13 months. Can be done in cash, short-term securities. Deposits, T-bills, CDs, commercial paper, OR repo agreements. There’s a contract to repurchase a security sometime in the future, which will correspond to a rate. There’s an interbank and eurocurrency markets. Banks are required to keep a percentage of deposits as reserve with their local Federal Reserve bank. Basically the banks will borrow reserves from each other overnight to avoid falling below this threshold. The effective federal funds rate is determined by the market, and influenced by the Federal Reserve, the open market operations. Also, there’s a reference rate, the SOFR (Secured OVernight FInancing Rate), this is a broad measure of the cost of borrowing cash overnight, which is collateralized by treasury securities. For a long time LIBOR was taken as a reference rate, SOFR has replaced it. SOFR is an evaluated median of transaction medium, repo data collected by the Bank of New York. The methodology is a little more complicated but this is a 50th percentile of volume. LIBOR was the rate at which banks were willing to loan to each other at. They didn’t necessarily have to have the transaction, they just had to post a quote. SOFR is data-driven. Many fixed income securities like government bonds, agency debt, municipal debt, etc. ### Creditworthiness Some interest rates we might consider to be risk-free, like treasuries or US government agency securities. There are also low-risk instruments, floating coupon bonds indexed by floating rate, swaps, futures, forward rate agreements. Here the idea is, by looking at the market, you’ll be able to find many such rates. So the yield curve you’ll find is not necessarily unique, here’s one for treasuries, corporate bonds, etc. Here we’re going to look at the construction of such yield curves. ## Bootstrapping Yield Curve ### Definitions A **coupon paying bond** is a contract that pays a fixed coupon at future times. Typically this is given in terms of the annual rate, like 5% per annum. Such coupon is based on the principal amount of the bond. This is paid with a certain frequency. This coupon bond has the cash flow of the reimbursement of the notional value of the bond. #### Example Let’s say we have a bond that pays 5% annual coupon with T \= 5 years until maturity. Let’s assume that the principal P \= $100. The cash flow diagram is very simple, with cash flows of $5 at each year after year 0, and then the max cash flow at the end. The **zero coupon bond** for a maturity t guarantees the payment of one unit at maturity. This is important because we use this extensively. The price of a zero coupon bond paying $1 at maturity t is called the **discount factor**, d(t) or P(t, T), the price of a bond maturing at T at time t. In general, the price of zero coupon is less than 1\. As the maturity of the contract increases, its value decreases. If you take $$\\frac{\\partial P(t, T)}{\\partial T} \< 0$$ A **par coupon bond** is worth par, 100% of the notional. This is c(t). A **forward rate agreement** is a commitment to lend money at a specified future rate for a specified period time. The rate of a forward loan is called forward rate. The forward rate is f(t) from t \- 1 to t. A **floating rate note** is a contract ensuring payment of a floating rate at future dates and pays a last cash flow reimbursing the notional at T. A **swap** pays fixed for floating to exchange periodic interest payment. ### Bond Pricing If you just observe the prices of the bonds, it doesn’t contain enough information to know value. You also have to compare the return holding the bond. One measure of the performance is the **Yield to Maturity**. This is the rate of return on a bond if it is held to maturity. This is expressed as an annual rate. It’s important to understand the frequency of the rate to compare fairly. Based on current price we can compute a discount rate such that the PV of the future bond cash flows match the current price. The constant discount rate represents the yield of the bond. #### Example Consider a coupon bond with 4 years to maturity at 10% annual coupon rate. Assume current bond price is $90. What is YTM? By definition, price is PV \= 90\. Typically you take the notional in these to be $100. Therefore the price is $$90 \= \\frac{10}{1+y} \+ \\frac{10}{(1+y)^2} \+ \\frac{10}{(1+y)^3} \+ \\frac{10}{(1+y)^4} \+ \\frac{100}{(1+y)^4}$$ Solve for y \= 0.1338 \= 13.38% This would be the yield of this bond. If the bond price is equal to the principal then what is the yield? $$100 \= \\frac{10}{1+y}...$$ Then y \= 0.10 \= 10%. So then the yield is the same as the coupon rate when the bond is par. ### Clean/Dirty Clean price, also known as quoted price, is the price of the bond excluding interest that has accrued since issue or the most recent coupon payment. So the price of the bond is supposed to include this accrual of the value of that bond. However, it is not quoted with this accrual, so the clean price is quoted. Otherwise you would have to adjust rates continuously. Dirty price, also known as cash price, is the price of a bond which includes the accrued interest. Let’s look at some calculations #### Example Consider buying a 3-year 12% annual coupon bond. (30/360 day count convention), 360 days and 30 days in a month, within one month from the first coupon. What is the YTM? The coupon value is c. Then the final cash flow is 100 \+ c. Let’s say you buy at time t which is before t1, one month before t1 the first coupon payment. There is some accrued interest at time t. The formula is the same, with the cash flows discounted to the time t. What is the accrual in this case? It’s of 11 months, which by our convention is 330/360 \* c. The coupon is 12% of the principal 100, therefore 12, so 330/360 \* 12 \= 11\. We can say as a general formula that $$Dirty Price \= Clean Price \+ Accrued Interest$$ $$100 \+ 11 \= \\frac{12}{(1+y)^{\\tfrac{30}{360}}} \+ \\frac{12}{(1+y)^{\\tfrac{390}{360}}} \+ \\frac{112}{(1+y)^{\\tfrac{750}{360}}}$$ Using a solver, you can find that y \= 11.97%. The quotations would have to adjust continuously. Dirty prices can be reported but they are not typically quoted. ### Sensitivity to Yield Consider a fixed income instrument with price P and yield Y. We will define DV01 as the dollar value of one basis point. $$DV01 \= \-\\frac{\\Delta P}{10,000 \\cdot \\Delta y}$$ Negative means that as one increases, the other decreases. We also have **duration**, which is $$D \= \-\\frac{1}{P} \\cdot \\frac{\\Delta P}{\\Delta y}$$ and **convexity** $$C \= \\frac{1}{P} \\cdot \\frac{\\Delta^2 P}{\\Delta y^2}$$ This is written as a difference, but you can generalize this as a derivative. You can approximate the rate of change of the price of a bond as being dependent on the duration and convexity with a formula combining them. #### Example 1 Consider a 1-year zero coupon bond with annual yield y. What is the relationship between P and y? $$P \= \\frac{1}{1+y}$$ Then using our formula for DV01, we get $$DV01 \= \-\\frac{1}{10,000} \\frac{dP}{dY}$$ We can use the quotient rule (or chain rule\!) to take the derivative, which results in $$= \\frac{1}{10,000(1+y)^2}$$ ## The Cost of Transit Delays - URL: https://sharifhsn.dev/blog/linkedin-2025-01-28-the-cost-of-transit-delays/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-01-28-the-cost-of-transit-delays.json - Description: As a frequent user of public transit, I know how impactful delays are both on my productivity and my mental health. Being stuck in Secaucus Junction because of some nonsense happen… - Date: 2025-01-28 - Exact published timestamp: 2025-01-28T14:46:42.541Z - Topics: Transportation, Public Policy - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7290017452040544257/ As a frequent user of public transit, I know how impactful delays are both on my productivity and my mental health. Being stuck in Secaucus Junction because of some nonsense happening in NYC is a feeling all too familiar for me. Researchers from Stevens Institute of Technology have quantified the impact of disinformation in the media by using AI to model disruption scenarios from the social media alerts for the PATH train. They used clustering techniques like K-means and BERTopic to classify different disruption scenarios, then used Monte Carlo simulation to generate potential train flows given these scenarios, calculating lost time and rerouting costs. They discovered that even mild disruptions could lead to stations being closed as much as 11% of the time, if timed at certain periods like major sporting events. I've seen how the waiting area fills up with disgruntled Knicks fans after a Hawks game, and it's a ticking time bomb for a targeted attack. In our modern, social media-driven world, the potential for information propagation causing mayhem is more significant than ever before. It is incumbent on us all to remain calm when such disruptions happen, knowing that the cause may not be mere negligence, but active maliciousness by an attacker. Read more here: [https://lnkd.in/epaDBfH5](https://lnkd.in/epaDBfH5) The underlying study: [https://lnkd.in/ehDt9RNv](https://lnkd.in/ehDt9RNv) ## Stochastic Processes and Computational Pricing - URL: https://sharifhsn.dev/blog/computational-methods-week-01/ - Structured data: https://sharifhsn.dev/api/posts/computational-methods-week-01.json - Description: I’m out on vacation this week, so I watched the recording - Date: 2025-01-28 - Exact published timestamp: 2025-01-28 - Topics: Computational Methods, Stochastic Processes, Numerical Methods, Option Pricing - Categories: Computational Methods - Source: Computational Methods in Quantitative Finance - Source URL: None I’m out on vacation this week, so I watched the recording ## Textbook ### Stochastic Processes A stochastic processes is merely a collection of indexed random variables. The indexing of these variables confers a structure onto the rvs which gives them some properties. The indexing structure can be a set or an interval, which would imply a discrete or continuous stochastic process, respectively. ## Syllabus Late assignments are not accepted under any circumstances, at least 24 hours in advance must be warned. Attendance is mandatory for in-person, not for online. Old book with pseudocode with formula. which is a lot more Main things we cover: Monte carlo approximation Finite difference Trees Black-Scholes PDE solution we study in 610 Most models do not have a formula though So how do we estimate the value then? That’s what this class is about There are two fundamental ways Approximate the process. Let’s say I know the path for sure. I can calculate the value by using the payoff formula. I can generate millions of paths and then average them. That is very slow. Or I could look at the probability of each path, and then average them? That’s trees. Or you could solve the PDE. You can get complicated PDEs, but they always have a very similar structure. ## Rust 1.84: Three Infrastructure Changes - URL: https://sharifhsn.dev/blog/linkedin-2025-01-10-rust-1-84-three-infrastructure-changes/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2025-01-10-rust-1-84-three-infrastructure-changes.json - Description: A new version of the Rust language has been released: Rust 1.84.0! - Date: 2025-01-10 - Exact published timestamp: 2025-01-10T13:25:09.372Z - Topics: Rust, Developer Tools - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7283473947021459456/ A new version of the Rust language has been released: Rust 1.84.0! Although there aren't any new language features, this release kickstarts three big infrastructure improvements for Rust that have been a long time coming: minimum supported Rust version (MSRV), a new trait solver, and pointer provenance. MSRV is a community-driven standard for Rust libraries that seek to be compatible with users that require older versions of Rust. For a while now, Rust has allowed libraries to declare an MSRV, but it doesn't do anything other than inform the user. If, for example, a library bumped the MSRV up when adding a new feature, the user would be forced to manually specify an older version of the library, which is hacky and difficult to maintain. Now, Rust's package manager Cargo will detect the user's MSRV and only download versions of libraries compatible with it. Rust's powerful type system uses type theory to evaluate the correctness of the code's type specifications. As new features have been requested from Rust over time, the system that solves these types has bolted on the required features as they are implemented, resulting in a hacky system that's difficult to extend. The new from-scratch trait solver is simpler and is a necessary prerequisite to some highly-requested features, like specialization. In systems programming languages like C, users will often need to work with raw pointers to memory i.e. integers. Rust uses references by default, which are pointers to memory that are tied to the object that created them, and therefore contain information relating to that object, like its mutability. This information is called provenance. References are what allows Rust to declare itself "memory-safe". However, it is sometimes necessary to construct pointers out of thin air, instead of references to an object, for example when working in an embedded environment. Previously, only raw pointers could be used for this, resulting in an unwieldy, buggy, C-like interface for users. Now, pointers can contain provenance, which gives them a safer API. Nothing flashy, but all great QOL improvements! [https://lnkd.in/eiBrWqPd](https://lnkd.in/eiBrWqPd) ## Exotic Options - URL: https://sharifhsn.dev/blog/exotic-options/ - Structured data: https://sharifhsn.dev/api/posts/exotic-options.json - Description: The final FE-610 notes apply the earlier stopping-time and maximum results to path-dependent payoffs. A barrier option depends on whether the underlying crosses a level before expi… - Date: 2024-12-12 - Exact published timestamp: 2024-12-12 - Topics: Stochastic Calculus, Exotic Options, Barrier Options - Categories: Stochastic Calculus - Source: FE-610 \| Stochastic Calculus - Source URL: None The final FE-610 notes apply the earlier stopping-time and maximum results to path-dependent payoffs. A barrier option depends on whether the underlying crosses a level before expiry; a lookback option depends on the running maximum or minimum. The running maximum \(M_t=\max_{0\leq u\leq t}W_u\) is not determined by the terminal value alone. The pair \((W_t,M_t)\) is the useful Markov state. Reflection arguments relate events involving the maximum to ordinary Brownian probabilities, which makes joint distributions and first-passage calculations possible. These products are a natural reason to study stochastic calculus: their value depends on the whole path, so a terminal-price shortcut is insufficient. ## Options and Volatility - URL: https://sharifhsn.dev/blog/options-and-volatility/ - Structured data: https://sharifhsn.dev/api/posts/options-and-volatility.json - Description: The final FE-535 notes turn to nonlinear risk through options. An option's payoff is convex, so a small change in the underlying can have a different effect depending on where the … - Date: 2024-12-12 - Exact published timestamp: 2024-12-12 - Topics: Risk Management, Options, Volatility - Categories: Risk Management - Source: FE-535 \| Risk Management - Source URL: None The final FE-535 notes turn to nonlinear risk through options. An option's payoff is convex, so a small change in the underlying can have a different effect depending on where the price sits relative to the strike. Delta is the local slope; gamma measures how quickly that slope changes. The VIX discussion is a reminder that implied volatility is a market price of uncertainty, not a direct forecast of realized volatility. Volatility can vary with strike and maturity, producing a surface rather than a single number. That curvature is why a delta hedge must be rebalanced. A position that looks neutral for one price move can acquire substantial exposure after the underlying moves or implied volatility changes. ## Default Fields in Rust Structs - URL: https://sharifhsn.dev/blog/linkedin-2024-12-08-default-fields-in-rust-structs/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2024-12-08-default-fields-in-rust-structs.json - Description: A recent RFC (request for comment) for the Rust language has divided the community: default fields in structs. - Date: 2024-12-08 - Exact published timestamp: 2024-12-08T16:50:39.228Z - Topics: Rust, Programming Languages - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7271566862621020161/ A recent RFC (request for comment) for the Rust language has divided the community: default fields in structs. In order to implement a default version of a struct in Rust, one has to implement the `Default` trait on it. The easiest way to do this is to use the derive macro for `Default`, which requires every single one of the structs fields to implement `Default`. This works fine for structs composed of simple elements that already implement `Default`, like integers and strings, but for more complex objects, you will need to either implement `Default` through an `impl` block for each of the struct's unimplemented fields, or have an `impl` block for the whole struct. Both of these workarounds have disadvantages. Implementing `Default` for every struct field type you design is tedious, especially if they are rarely used in a `Default` context. Additionally, it is often possible that a struct field type needs different `Default` implementations in different parent struct contexts, rendering this workaround impossible. The other solution, writing a bespoke `impl Default` block for every parent struct, is prone to code duplication, a common source of errors. The simple solution is default field types. These allow users to specify the default initializer for a struct field type tersely within a parent struct. You can see in this attached image how much easier this is. The controversy arises from the fact that Rust already has ways of dealing with this problem. In addition to the workarounds I mentioned earlier, the "pure" way of doing this would involve instantiating a new struct field type for every different `Default` implementation. In the example I've shown here, instead of being an `i128`, the `age` field would have a special type `PetAge(i128)` which would have a separate `impl Default`. However, this creates a lot of boilerplate and bloats code size, which is always a code smell. Additionally, the RFC has the limitation that the in-struct `Default` initializer must be `const`, a limitation that `impl Default` *does not face*. This is important because some of the most common use-cases of this in third-party derive macros that provide this functionality are for non-`const` initialization, particularly `String` initialization. You can read more discussion about this on the RFC pull request page: [https://lnkd.in/eTZ8FPd7](https://lnkd.in/eTZ8FPd7) What do you think of this upcoming feature? Is it useful, or just needless complexity? ## Hedging Bonds - URL: https://sharifhsn.dev/blog/bond-hedging/ - Structured data: https://sharifhsn.dev/api/posts/bond-hedging.json - Description: The bond-hedging lab combines duration, futures, and forward positions. A portfolio can be made locally insensitive to a yield move by matching its dollar duration with an offsetti… - Date: 2024-12-05 - Exact published timestamp: 2024-12-05 - Topics: Risk Management, Bonds, Hedging - Categories: Risk Management - Source: FE-535 \| Risk Management - Source URL: None The bond-hedging lab combines duration, futures, and forward positions. A portfolio can be made locally insensitive to a yield move by matching its dollar duration with an offsetting instrument. The hedge has to match the exposure being measured. A bond's price, accrued interest, maturity, coupon, and day-count convention determine the sensitivity; a futures contract adds its own conversion factor and basis. Matching only the face amount can leave a large residual rate exposure. The notes use this as a practical version of the earlier lesson: a hedge is a model of a risk, and the model must be re-estimated as the portfolio and the curve change. ## Rust 1.83 Brings More `const` - URL: https://sharifhsn.dev/blog/linkedin-2024-11-28-rust-1-83-brings-more-const/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2024-11-28-rust-1-83-brings-more-const.json - Description: A happy Thanksgiving to all! Today, in addition to my loved ones, I'm grateful for a brand-new version of Rust, now with more const! const evaluation is one of the most important s… - Date: 2024-11-28 - Exact published timestamp: 2024-11-28T16:16:30.611Z - Topics: Rust, Programming Languages - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/feed/update/urn:li:activity:7267934391442583554/ A happy Thanksgiving to all! Today, in addition to my loved ones, I'm grateful for a brand-new version of Rust, now with more `const`! `const` evaluation is one of the most important steps in making Rust more ergonomic in memory-constrained environments, such as embedded computing or the Linux kernel, as it allows for more code to be evaluated before execution. For example, let's say that a program requires the result of some complex mathematical expression to run. Computing this at runtime would add significant overhead and could violate memory constraints. By placing this computation in a `const` context (either a `const` expression or a `const` function), the computation is guaranteed to run at time of compilation, with the result stored in the program to be used immediately upon execution. This new release allows users to take mutable references in a `const` context, greatly expanding what kind of expressions are allowed and improving the flexibility of `const` code. As a result, many common functions in the standard library are now stabilized in the `const` context, including many operations on floating point numbers. I'm looking forward to how this new feature will enable better patterns in embedded Rust! [https://lnkd.in/eC3x_U52](https://lnkd.in/eC3x_U52) ## Poisson Processes - URL: https://sharifhsn.dev/blog/poisson-processes/ - Structured data: https://sharifhsn.dev/api/posts/poisson-processes.json - Description: The FE-610 notes introduce Poisson processes as counting processes with independent increments and a constant arrival intensity. For a rate \(\lambda\), the number of arrivals by t… - Date: 2024-11-28 - Exact published timestamp: 2024-11-28 - Topics: Stochastic Calculus, Poisson Processes, Jump Diffusion - Categories: Stochastic Calculus - Source: FE-610 \| Stochastic Calculus - Source URL: None The FE-610 notes introduce Poisson processes as counting processes with independent increments and a constant arrival intensity. For a rate \(\lambda\), the number of arrivals by time \(t\) has distribution $$\mathbb P(N_t=k)=e^{-\lambda t}\frac{(\lambda t)^k}{k!}.$$ The waiting time to the first arrival is exponential, and independent waiting times produce the full process. Adding random jump sizes gives a compound Poisson process; combining it with Brownian motion gives a jump-diffusion model. The notes return to quadratic variation: jumps contribute their squared sizes, while the continuous part contributes its usual diffusion variation. That decomposition is the starting point for an Itô formula with jumps. ## Linear Hedging - URL: https://sharifhsn.dev/blog/linear-hedging/ - Structured data: https://sharifhsn.dev/api/posts/linear-hedging.json - Description: The linear-risk notes approximate a portfolio's change by its sensitivities to risk factors. A hedge chooses positions whose factor exposures offset the portfolio's exposures, ofte… - Date: 2024-11-21 - Exact published timestamp: 2024-11-21 - Topics: Risk Management, Hedging, Linear Risk - Categories: Risk Management - Source: FE-535 \| Risk Management - Source URL: None The linear-risk notes approximate a portfolio's change by its sensitivities to risk factors. A hedge chooses positions whose factor exposures offset the portfolio's exposures, often by solving a covariance-weighted least-squares problem. For a single factor, the minimum-variance hedge ratio is proportional to covariance divided by the variance of the hedging instrument. The unitary hedge examples show why the sign and units must be checked: a hedge ratio is a position size, not a probability. The approximation is local. Basis risk, nonlinear payoffs, liquidity, and changing correlations can all make the realized hedge error larger than the linear estimate. ## Stochastic Differential Equations - URL: https://sharifhsn.dev/blog/stochastic-differential-equations/ - Structured data: https://sharifhsn.dev/api/posts/stochastic-differential-equations.json - Description: The SDE section writes a state process as a drift plus a diffusion term, - Date: 2024-11-21 - Exact published timestamp: 2024-11-21 - Topics: Stochastic Calculus, SDEs, Diffusions - Categories: Stochastic Calculus - Source: FE-610 \| Stochastic Calculus - Source URL: None The SDE section writes a state process as a drift plus a diffusion term, $$dX_t=b(t,X_t)\,dt+\sigma(t,X_t)\,dW_t.$$ The coefficients describe deterministic motion and random shocks. Unlike an ordinary differential equation, an SDE is interpreted through an integral equation and a chosen filtration. The notes compare continuous diffusions with jump processes and use Itô's formula to transform solutions. For pricing, this notation is valuable because a model can be specified by its local characteristics even when there is no closed-form path. Existence, integrability, and the chosen measure determine whether the process is usable as a financial model. ## Multidimensional Market Models - URL: https://sharifhsn.dev/blog/multidimensional-market-models/ - Structured data: https://sharifhsn.dev/api/posts/multidimensional-market-models.json - Description: The multidimensional market model in the FE-610 notes has several stocks driven by several Brownian factors. Prices, drifts, volatilities, and correlations become vectors and matri… - Date: 2024-11-14 - Exact published timestamp: 2024-11-14 - Topics: Stochastic Calculus, Market Models, Girsanov - Categories: Stochastic Calculus - Source: FE-610 \| Stochastic Calculus - Source URL: None The multidimensional market model in the FE-610 notes has several stocks driven by several Brownian factors. Prices, drifts, volatilities, and correlations become vectors and matrices, but the no-arbitrage idea is unchanged: after discounting, tradable prices should be martingales under an equivalent measure. The multidimensional Girsanov change shifts the Brownian vector by a market-price-of-risk process. The fundamental theorem then links existence of an equivalent martingale measure to no arbitrage, and uniqueness to completeness. In practice, the model's volatility matrix determines whether every contingent claim can be replicated. ## Risk-Neutral Valuation and Interest Rate Parity - URL: https://sharifhsn.dev/blog/risk-neutral-valuation-and-parity/ - Structured data: https://sharifhsn.dev/api/posts/risk-neutral-valuation-and-parity.json - Description: The risk-neutral valuation notes use a pricing measure under which discounted tradable prices are martingales. The expected payoff is discounted at the funding rate rather than for… - Date: 2024-11-14 - Exact published timestamp: 2024-11-14 - Topics: Risk Management, Risk-Neutral Valuation, Interest Rate Parity - Categories: Risk Management - Source: FE-535 \| Risk Management - Source URL: None The risk-neutral valuation notes use a pricing measure under which discounted tradable prices are martingales. The expected payoff is discounted at the funding rate rather than forecast under the physical return distribution. For foreign exchange, the same replication logic gives interest-rate parity. If domestic and foreign money-market accounts are both available, the forward exchange rate must balance their growth rates; otherwise borrowing in one currency and lending in the other creates an arbitrage. The class treats parity as a control check. A quoted forward, spot, and pair of rates should agree after day-count, compounding, and collateral conventions are made explicit. ## Joint and Conditional Distributions - URL: https://sharifhsn.dev/blog/joint-and-conditional-distributions/ - Structured data: https://sharifhsn.dev/api/posts/joint-and-conditional-distributions.json - Description: The final FE-540 notes work with a joint law rather than treating each variable in isolation. A conditional density is obtained by normalizing the joint density with the relevant m… - Date: 2024-11-11 - Exact published timestamp: 2024-11-11 - Topics: Probability Theory, Conditional Distributions, Covariance - Categories: Probability Theory - Source: FE-540 \| Probability Theory - Source URL: None The final FE-540 notes work with a joint law rather than treating each variable in isolation. A conditional density is obtained by normalizing the joint density with the relevant marginal: $$f_{X\mid Y}(x\mid y)=\frac{f_{X,Y}(x,y)}{f_Y(y)}.$$ This makes conditional expectation a function of the information being observed. Covariance records co-movement, \(\operatorname{Cov}(X,Y)=\mathbb E[(X-\mathbb E X)(Y-\mathbb E Y)]\), while conditional versions let the same idea change as information arrives. The notes' worked examples are a useful reminder to state the support first, integrate over the correct region, and check that the resulting conditional density integrates to one. ## Forward Contracts - URL: https://sharifhsn.dev/blog/forward-contracts/ - Structured data: https://sharifhsn.dev/api/posts/forward-contracts.json - Description: The forward-contract notes build value from replication. A forward fixes the delivery price today, while the underlying can be financed or invested until maturity. The forward pric… - Date: 2024-11-07 - Exact published timestamp: 2024-11-07 - Topics: Risk Management, Forward Contracts, Futures - Categories: Risk Management - Source: FE-535 \| Risk Management - Source URL: None The forward-contract notes build value from replication. A forward fixes the delivery price today, while the underlying can be financed or invested until maturity. The forward price is therefore tied to the spot price, financing rate, and any income or storage benefit from holding the underlying. For an asset with a known cash yield, the no-arbitrage relation has the form $$F_0=S_0e^{(r-q)T},$$ with the appropriate convention for the yield or income rate. Futures add daily settlement and margin, so their value can differ from a forward when rates and prices are correlated. The risk-management use is direct: a forward can lock a future price, but it replaces price uncertainty with counterparty, funding, and liquidity exposures. ## Risk-Neutral Measures - URL: https://sharifhsn.dev/blog/risk-neutral-measures/ - Structured data: https://sharifhsn.dev/api/posts/risk-neutral-measures.json - Description: The risk-neutral section motivates a change of measure rather than assuming investors literally expect the risk-free rate. Discounted tradable prices should be martingales under a … - Date: 2024-11-07 - Exact published timestamp: 2024-11-07 - Topics: Stochastic Calculus, Risk-Neutral Pricing, Girsanov - Categories: Stochastic Calculus - Source: FE-610 \| Stochastic Calculus - Source URL: None The risk-neutral section motivates a change of measure rather than assuming investors literally expect the risk-free rate. Discounted tradable prices should be martingales under a pricing measure \(\mathbb Q\). Under that measure, a payoff \(H_T\) is valued as $$V_t=\mathbb E^{\mathbb Q}\left[e^{-r(T-t)}H_T\mid\mathcal F_t\right].$$ Girsanov's theorem explains how the drift changes when the Brownian motion is shifted. The new measure is equivalent to the original one, so zero-probability events stay zero; only the weights of possible paths change. The notes use the density process and the discounted asset to connect this measure change to no-arbitrage pricing. ## A Rainbow on Election Day - URL: https://sharifhsn.dev/blog/linkedin-2024-11-06-a-rainbow-on-election-day/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2024-11-06-a-rainbow-on-election-day.json - Description: A brief pause from election-day politics to notice an unexpected rainbow. - Date: 2024-11-06 - Exact published timestamp: 2024-11-06T19:56:37.382Z - Topics: Campus Life, Personal - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/posts/sharif-haason_i-was-walking-on-campus-this-morning-preoccupied-activity-7260017251611815936-2p_J I was walking on campus this morning, preoccupied with thoughts of the election, when I noticed a sprinkler on some flowers next to me. The water caught the light in just the right way to create a huge rainbow. The beauty of it stunned me, and I stood for a minute just staring at it. I reflected on how the small things in life are what's important. There will be plenty of time to debate and discuss the impact of this election, but that rainbow will only exist for a brief period. If your thoughts, like mine, are stormy with politics today, take a moment and look for a rainbow near you. ## Random Vectors and Independence - URL: https://sharifhsn.dev/blog/random-vectors-and-independence/ - Structured data: https://sharifhsn.dev/api/posts/random-vectors-and-independence.json - Description: The later FE-540 notes package several random variables into a random vector, such as \(X=(X_1,X_2)\). The joint distribution records probability over rectangles and more general s… - Date: 2024-11-04 - Exact published timestamp: 2024-11-04 - Topics: Probability Theory, Random Vectors, Independence - Categories: Probability Theory - Source: FE-540 \| Probability Theory - Source URL: None The later FE-540 notes package several random variables into a random vector, such as \(X=(X_1,X_2)\). The joint distribution records probability over rectangles and more general subsets of the product space. Marginal distributions come from integrating or summing out the other coordinates. Independence has a clean joint-density form when densities exist: $$f_{X,Y}(x,y)=f_X(x)f_Y(y).$$ The same factorization can be stated with a joint CDF or with sigma-algebras. It is stronger than zero covariance: independent variables have zero covariance when the moments exist, but uncorrelated variables need not be independent. The notes use this distinction before moving into conditional distributions and covariance calculations. ## Derivatives and Clearinghouses - URL: https://sharifhsn.dev/blog/derivatives-and-clearinghouses/ - Structured data: https://sharifhsn.dev/api/posts/derivatives-and-clearinghouses.json - Description: The derivatives notes frame a contract as a way to move or reshape risk. Futures and standardized options trade through a clearinghouse, while bilateral contracts expose each party… - Date: 2024-10-31 - Exact published timestamp: 2024-10-31 - Topics: Risk Management, Derivatives, Clearinghouses - Categories: Risk Management - Source: FE-535 \| Risk Management - Source URL: None The derivatives notes frame a contract as a way to move or reshape risk. Futures and standardized options trade through a clearinghouse, while bilateral contracts expose each party to counterparty risk. Central counterparties reduce bilateral connections by becoming the buyer to every seller and the seller to every buyer. That changes the network of exposures; it does not make risk disappear. Margin, default funds, collateral, and close-out rules determine how losses are absorbed when a member fails. The class uses the Robinhood and clearinghouse discussion to connect market structure with risk measurement. A position's payoff, collateral terms, liquidity, and legal netting set all belong in the exposure analysis. ## Stochastic Calculus Review - URL: https://sharifhsn.dev/blog/stochastic-calculus-review/ - Structured data: https://sharifhsn.dev/api/posts/stochastic-calculus-review.json - Description: The FE-610 review consolidates the first half of the course: probability spaces and filtrations, martingales, Brownian motion, quadratic variation, the Itô integral, Itô's formula,… - Date: 2024-10-31 - Exact published timestamp: 2024-10-31 - Topics: Stochastic Calculus, Review, Pricing - Categories: Stochastic Calculus - Source: FE-610 \| Stochastic Calculus - Source URL: None The FE-610 review consolidates the first half of the course: probability spaces and filtrations, martingales, Brownian motion, quadratic variation, the Itô integral, Itô's formula, and the Black–Scholes hedge. The compact rules worth carrying forward are $$dW_t\,dt=0,\qquad (dW_t)^2=dt,\qquad (dt)^2=0.$$ They are a mnemonic for the limiting variation calculation, not ordinary algebra. The review also separates three questions that are easy to conflate: what is measurable now, which process is a martingale under the current measure, and which payoff is being priced. ## Student's t and Heavy-Tailed Distributions - URL: https://sharifhsn.dev/blog/student-t-and-heavy-tailed-distributions/ - Structured data: https://sharifhsn.dev/api/posts/student-t-and-heavy-tailed-distributions.json - Description: After the gamma and beta families, the FE-540 notes turn to Student's \(t\) distribution and other shapes used when observations are more extreme than a Gaussian model suggests. St… - Date: 2024-10-28 - Exact published timestamp: 2024-10-28 - Topics: Probability Theory, Student t, Heavy Tails - Categories: Probability Theory - Source: FE-540 \| Probability Theory - Source URL: None After the gamma and beta families, the FE-540 notes turn to Student's \(t\) distribution and other shapes used when observations are more extreme than a Gaussian model suggests. Student's \(t\) can be built from a standard normal variable divided by the square root of an independent chi-squared variable scaled by its degrees of freedom. The notes also compare Pareto, log-normal, and Laplace distributions. The point is not just to collect names: the tail behavior controls whether moments exist and how sensitive an estimate is to a few unusually large observations. A distribution with a finite mean can still have a very unstable variance, while a log-normal or Pareto tail can make sample averages converge slowly. Writing down the support and checking the normalizing constant are the reliable first steps before calculating a moment. ## Remembering the Boarding-School History - URL: https://sharifhsn.dev/blog/linkedin-2024-10-27-remembering-the-boarding-school-history/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2024-10-27-remembering-the-boarding-school-history.json - Description: A reflection on President Biden's apology for federal Indian boarding schools. - Date: 2024-10-27 - Exact published timestamp: 2024-10-27T15:20:30.081Z - Topics: Education, Native American History - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/posts/sharif-haason_president-biden-apologizes-for-treatment-activity-7256323884658888704-1-gA I was moved by this emotional announcement by President Biden, apologizing for an often forgotten aspect of oppression of Native Americans: the forced kidnapping of their children to boarding schools where they were stripped of their language and culture. This policy was only ended in 1978 with the Indian Child Welfare Act. It's a common attitude of Americans that the mistreatment of Native Americans was something that happened a long time ago. Statements like these by our president are therefore important for educating the public and I urge you to listen to his speech and understand the magnitude of suffering that the history books have skipped over. ## Interest Rate Risk and Duration - URL: https://sharifhsn.dev/blog/interest-rate-risk-and-duration/ - Structured data: https://sharifhsn.dev/api/posts/interest-rate-risk-and-duration.json - Description: The interest-rate-risk notes approximate how a bond portfolio changes when yields move. Duration is the first-order sensitivity of price to yield; convexity captures the curvature … - Date: 2024-10-24 - Exact published timestamp: 2024-10-24 - Topics: Risk Management, Interest Rate Risk, Duration - Categories: Risk Management - Source: FE-535 \| Risk Management - Source URL: None The interest-rate-risk notes approximate how a bond portfolio changes when yields move. Duration is the first-order sensitivity of price to yield; convexity captures the curvature that duration misses. For a price \(P(y)\), the class defines modified duration and DV01 through the local derivative: $$D=-\frac{1}{P}\frac{\partial P}{\partial y},\qquad DV01=-\frac{\partial P}{10{,}000\,\partial y}.$$ The signs reflect the inverse relation between price and yield. A duration-only hedge is adequate for a small parallel move, while convexity matters for larger moves or portfolios whose cash flows are spread across maturities. ## Multidimensional Itô Calculus - URL: https://sharifhsn.dev/blog/multidimensional-ito-calculus/ - Structured data: https://sharifhsn.dev/api/posts/multidimensional-ito-calculus.json - Description: When several state variables move together, the FE-610 notes replace the scalar Itô formula with its multidimensional version. The Hessian term contains both variances and cross-va… - Date: 2024-10-24 - Exact published timestamp: 2024-10-24 - Topics: Stochastic Calculus, Multidimensional Models, Correlation - Categories: Stochastic Calculus - Source: FE-610 \| Stochastic Calculus - Source URL: None When several state variables move together, the FE-610 notes replace the scalar Itô formula with its multidimensional version. The Hessian term contains both variances and cross-variations, so correlation enters the drift of a function of two stocks. For a two-dimensional process driven by correlated Brownian motions, the covariance matrix is part of the model. A portfolio or derivative can therefore depend on both individual volatilities and the covariance term. The notes use a two-dimensional Itô process and Lévy's characterization to keep track of these cross terms. This is the bridge from one-stock Black–Scholes to a market model with several assets: the same calculus works, but the matrix of quadratic covariations must be carried through every derivative. ## Gamma, Beta, and Chi-Squared Distributions - URL: https://sharifhsn.dev/blog/gamma-beta-and-chi-squared-distributions/ - Structured data: https://sharifhsn.dev/api/posts/gamma-beta-and-chi-squared-distributions.json - Description: The FE-540 distribution notes connect several continuous families through the gamma function. The chi-squared family is a gamma distribution with parameters \(n/2\) and \(1/2\), wh… - Date: 2024-10-21 - Exact published timestamp: 2024-10-21 - Topics: Probability Theory, Gamma Distribution, Beta Distribution - Categories: Probability Theory - Source: FE-540 \| Probability Theory - Source URL: None The FE-540 distribution notes connect several continuous families through the gamma function. The chi-squared family is a gamma distribution with parameters \(n/2\) and \(1/2\), which makes its moments and additivity properties easier to remember. The beta function normalizes densities on \([0,1]\): $$\mathrm B(a,b)=\int_0^1x^{a-1}(1-x)^{b-1}\,dx =\frac{\Gamma(a)\Gamma(b)}{\Gamma(a+b)}.$$ For \(X\sim\operatorname{Beta}(a,b)\), the notes derive \(\mathbb E[X]=a/(a+b)\) and \(\operatorname{Var}(X)=ab/((a+b)^2(a+b+1))\). These identities are useful because they turn repeated integrals into parameter substitutions. ## Black–Scholes - URL: https://sharifhsn.dev/blog/black-scholes/ - Structured data: https://sharifhsn.dev/api/posts/black-scholes.json - Description: The Black–Scholes notes start with a stock following geometric Brownian motion, - Date: 2024-10-17 - Exact published timestamp: 2024-10-17 - Topics: Stochastic Calculus, Black–Scholes, Option Pricing - Categories: Stochastic Calculus - Source: FE-610 \| Stochastic Calculus - Source URL: None The Black–Scholes notes start with a stock following geometric Brownian motion, $$dS_t=\mu S_t\,dt+\sigma S_t\,dW_t.$$ Constructing a delta-hedged portfolio removes the Brownian shock. Applying Itô's formula to the option value and matching the remaining drift produces the Black–Scholes PDE. With terminal payoff \(g(S_T)\), the problem is a boundary-value problem: the price is determined backward from the payoff. The notes work through calls, puts, and the Black–Scholes–Merton form with a continuously compounded rate. The important modeling move is the hedge: risk is removed locally, so the option's drift is tied to the financing rate rather than to the stock's expected return. ## Probability and Regression Review - URL: https://sharifhsn.dev/blog/probability-and-regression-review/ - Structured data: https://sharifhsn.dev/api/posts/probability-and-regression-review.json - Description: The first-exam review returns to probability, conditional expectation, and regression. A risk estimate is a conditional statement: it depends on the information set, the horizon, a… - Date: 2024-10-17 - Exact published timestamp: 2024-10-17 - Topics: Risk Management, Probability, Regression - Categories: Risk Management - Source: FE-535 \| Risk Management - Source URL: None The first-exam review returns to probability, conditional expectation, and regression. A risk estimate is a conditional statement: it depends on the information set, the horizon, and the event being measured. Regression makes that dependence explicit. A market beta is a slope from a chosen sample and factor, not a permanent property of the asset. Residual risk remains after the factor is removed, and parameter uncertainty can be material when the sample is short or the regime has changed. The notes use this review to connect probability calculations with practical risk measurement: define the random variable, state the conditioning information, and check whether the residual assumptions are plausible. ## Continuous Random Variables and Transformations - URL: https://sharifhsn.dev/blog/continuous-random-variables-and-transformations/ - Structured data: https://sharifhsn.dev/api/posts/continuous-random-variables-and-transformations.json - Description: The continuous part of the FE-540 notes replaces sums with integrals. If \(X\) has density \(f_X\), then - Date: 2024-10-14 - Exact published timestamp: 2024-10-14 - Topics: Probability Theory, Continuous Distributions, Change of Variables - Categories: Probability Theory - Source: FE-540 \| Probability Theory - Source URL: None The continuous part of the FE-540 notes replaces sums with integrals. If \(X\) has density \(f_X\), then $$\mathbb E[h(X)]=\int_{-\infty}^{\infty}h(x)f_X(x)\,dx.$$ The same idea gives probabilities by integrating the density over an interval. The notes then work through a monotone transformation \(Y=h(X)\). For an invertible differentiable map, the density picks up the Jacobian factor: $$f_Y(y)=f_X\left(h^{-1}(y)\right)\left|\frac{d}{dy}h^{-1}(y)\right|.$$ The absolute derivative is the part that preserves mass when the coordinate system changes. The derivation proceeds through the CDF, a change of variables, and differentiation, which is safer than memorizing the final expression without its support. ## Bonds and the Time Value of Money - URL: https://sharifhsn.dev/blog/bonds-and-time-value-of-money/ - Structured data: https://sharifhsn.dev/api/posts/bonds-and-time-value-of-money.json - Description: The bond material in the FE-535 notes treats a fixed-coupon bond as a stream of future cash flows. Discounting those cash flows gives its present value, and the yield is the consta… - Date: 2024-10-10 - Exact published timestamp: 2024-10-10 - Topics: Risk Management, Bonds, Time Value of Money - Categories: Risk Management - Source: FE-535 \| Risk Management - Source URL: None The bond material in the FE-535 notes treats a fixed-coupon bond as a stream of future cash flows. Discounting those cash flows gives its present value, and the yield is the constant rate that makes the present value equal to the quoted price. The notes distinguish clean price from dirty price. Accrued interest belongs in the cash price even though market quotations commonly show the clean price: $$\text{Dirty Price}=\text{Clean Price}+\text{Accrued Interest}.$$ This decomposition is operationally important for a risk report. Two prices can differ simply because one includes the coupon accrued since the last payment. Day-count convention and payment frequency must be stated before comparing yields. ## Itô Calculus - URL: https://sharifhsn.dev/blog/ito-calculus/ - Structured data: https://sharifhsn.dev/api/posts/ito-calculus.json - Description: Itô's formula is the stochastic analogue of the multivariable chain rule. If - Date: 2024-10-10 - Exact published timestamp: 2024-10-10 - Topics: Stochastic Calculus, Ito Lemma, Geometric Brownian Motion - Categories: Stochastic Calculus - Source: FE-610 \| Stochastic Calculus - Source URL: None Itô's formula is the stochastic analogue of the multivariable chain rule. If $$dX_t=a(t,X_t)\,dt+b(t,X_t)\,dW_t,$$ then a smooth function \(f(t,X_t)\) acquires a second-order term because \((dW_t)^2=dt\): $$df=f_t\,dt+f_x\,dX_t+\tfrac12 f_{xx}(dX_t)^2.$$ The FE-610 notes apply the formula to Itô processes, generalized geometric Brownian motion, and short-rate examples such as Vasicek and Cox–Ingersoll–Ross. The extra curvature term is the practical difference from ordinary calculus; dropping it produces the wrong drift for a transformed process. ## Minneapolis Housing and Zoning - URL: https://sharifhsn.dev/blog/linkedin-2024-10-08-minneapolis-housing-and-zoning/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2024-10-08-minneapolis-housing-and-zoning.json - Description: What Minneapolis changed in its zoning rules, and what happened to housing and rents. - Date: 2024-10-08 - Exact published timestamp: 2024-10-08T21:50:44.977Z - Topics: Housing, Urban Policy - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/posts/sharif-haason_minneapolis-land-use-reforms-offer-a-blueprint-activity-7249536724404432898-dpDs Some of you likely tuned into the vice presidential debate last week. Tim Walz touted the success of housing policy in Minneapolis, so what did that actually look like? And how can those lessons be applied to the nationwide housing problem? The key behind Minneapolis’s housing was the Minneapolis 2040 Plan, approved in 2019 and put into place starting in 2020. This was a comprehensive set of 100 policies designed towards the goals of making housing affordable, increasing jobs, and making sure that urban areas fulfill environmental and social considerations. As all developers know, the primary obstacle to building housing is zoning. If a particular area is zoned to only allow single-family homes (SFHs), then the amount of development that can occur in that area is limited, as SFHs are large relative to their number of residents. There’s a fundamental density limitation which prevents SFHs from solving the housing problem in cities. The 2040 Plan upzoned Minneapolis in two ways. First, it requires residential areas which were SFH zoned to allow the building of duplexes and triplexes. This policy had very little impact on actual development because the rules required these homes to have the same floor-to-area ratio (FAR) as SFHs to prevent them from being larger. This placed significant design constraints on duplexes/triplexes to the point that they weren’t worth developing. A much more successful initiative was multifamily zoning on through streets, which the 2040 Plan calls “corridors”. These streets have bustled with business activity for ages, but their growth was stunted by their SFH zoning. The FAR requirements for corridor housing were relaxed, which allowed larger apartment buildings with 20-50 housing units to be constructed. This type of construction was the vast majority of housing that was permitted and built. There’s one more deregulation that developers say has been instrumental here: the elimination of citywide parking minimums. Minneapolis, like many other cities, used to require apartments in some areas to have parking. This meant that developers had to build additional parking structures, which studies show adds about $50,000 in per-unit development costs. It’s no wonder growth was stunted with such a massive burden on developers. What was the impact of this? Between 2017 and 2022, rents in Minneapolis stayed the same, when rents in Minnesota in general rose 14%. The levels of homelessness over the same period dropped 12% in Minneapolis and rose 14% in Minnesota. High rent and homelessness are the most pressing issues in the U.S. related to housing policy, so the plan can be considered a resounding success and a model for similar cities. What do you think of these policies? Should they be implemented in your city? Read the linked analysis: [Minneapolis Land Use Reforms Offer a Blueprint for Housing Affordability](https://www.pewtrusts.org/en/research-and-analysis/articles/2024/01/04/minneapolis-land-use-reforms-offer-a-blueprint-for-housing-affordability). ## Stochastic Integrals - URL: https://sharifhsn.dev/blog/stochastic-integrals/ - Structured data: https://sharifhsn.dev/api/posts/stochastic-integrals.json - Description: The FE-610 notes define an integral such as \(\int_0^T\Delta_t\,dW_t\) by starting with simple adapted processes. On each partition interval, the position \(\Delta_t\) is chosen fr… - Date: 2024-10-03 - Exact published timestamp: 2024-10-03 - Topics: Stochastic Calculus, Stochastic Integrals, Ito Integral - Categories: Stochastic Calculus - Source: FE-610 \| Stochastic Calculus - Source URL: None The FE-610 notes define an integral such as \(\int_0^T\Delta_t\,dW_t\) by starting with simple adapted processes. On each partition interval, the position \(\Delta_t\) is chosen from information already available, while the increment comes from Brownian motion. Refining the partition gives the Itô integral. The adaptedness condition is the key financial interpretation: a trading strategy can react to the past but cannot see the next Brownian increment. The integral is itself a random variable because every Brownian path produces a different gain. Two results organize the construction. The Itô integral is a martingale under suitable integrability, and the Itô isometry relates its second moment to the ordinary time integral of the squared integrand: $$\mathbb E\left[\left(\int_0^T\Delta_t\,dW_t\right)^2\right]=\mathbb E\left[\int_0^T\Delta_t^2\,dt\right].$$ ## A CUPS Vulnerability Explained - URL: https://sharifhsn.dev/blog/linkedin-2024-09-27-cups-linux-vulnerability/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2024-09-27-cups-linux-vulnerability.json - Description: An overview of the 2024 CUPS printer-discovery vulnerability and its risk. - Date: 2024-09-27 - Exact published timestamp: 2024-09-27T11:29:52.298Z - Topics: Cybersecurity, Linux - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/posts/sharif-haason_attacking-unix-systems-via-cups-part-i-activity-7245394208822300672-9IQ7 A 9.9 Linux CVE was discovered! Heartbleed was only 8.5, so you know that this is something to pay attention to. If you know any server administrators, text them some encouragement as they’ll surely be going through it at work today. Here’s some more detail on what the exploit is: ⬇️ CUPS is the Common Unix Printing System, a general service for printers that is used widely by Unix-like systems, including most distributions of Linux and macOS. The vulnerability was found in its printer discovery service `cups-browsed`. CUPS listens for packets at port 631 from *any host*. These packets trigger printer discovery, where CUPS tries to get the printer attributes from this arbitrary server. It can set its own name and the URL for the printing protocol to access, which could be an attacker-controlled host. There is also a printer attribute called `FoomaticRIPCommandLine`, which *arbitrarily executes* its string value when a print job is executed. This exists to provide flexibility for systems that need configuration before they can print… but it’s also a huge attack vector and has been the site of multiple previous CVEs. In essence: any computer connected to the public Internet is vulnerable to an attacker adding a malicious printer to it. If a user executes a print job with the malicious printer, then the arbitrary code, which could be malware, is executed. Keep in mind that the attacker can name the printer whatever they like. How many of you would think twice before printing to “HP DeskJet”? How is this mitigated? The easiest way to do this is to disable the `cups-browsed` service. Automatic printer discovery is more of a nice-to-have than a need, and with such a glaring security flaw, the correct decision is obvious. In general, considering how flawed CUPS is, you should avoid printing from Unix-like systems when possible. Want to know more? The author of the CVE published an excellent article that goes over all of the details of the exploit: [Attacking UNIX Systems via CUPS, Part I](https://www.evilsocket.net/2024/09/26/Attacking-UNIX-systems-via-CUPS-Part-I/). ## Brownian Motion - URL: https://sharifhsn.dev/blog/brownian-motion/ - Structured data: https://sharifhsn.dev/api/posts/brownian-motion.json - Description: Brownian motion is the continuous limit of a scaled symmetric random walk. Its increments are independent and normally distributed, with \(W_t-W_s\sim N(0,t-s)\). It is a martingal… - Date: 2024-09-26 - Exact published timestamp: 2024-09-26 - Topics: Stochastic Calculus, Brownian Motion, Quadratic Variation - Categories: Stochastic Calculus - Source: FE-610 \| Stochastic Calculus - Source URL: None Brownian motion is the continuous limit of a scaled symmetric random walk. Its increments are independent and normally distributed, with \(W_t-W_s\sim N(0,t-s)\). It is a martingale, but its paths are almost surely continuous and nowhere differentiable. The defining calculation for stochastic calculus is its quadratic variation: $$[W,W]_t=t,$$ or, in differential notation, \((dW_t)^2=dt\). Cross variation with ordinary time is zero. The notes use this contrast to explain why the ordinary chain rule cannot simply be applied to a Brownian path. First-passage times and the running maximum appear here as well. They connect Brownian motion to barrier and lookback payoffs later in the course. ## Brownian Motion and Geometric Growth - URL: https://sharifhsn.dev/blog/brownian-motion-risk-models/ - Structured data: https://sharifhsn.dev/api/posts/brownian-motion-risk-models.json - Description: The stochastic-process notes move from a random walk to Brownian motion and then to geometric Brownian motion. Brownian increments are normally distributed with variance proportion… - Date: 2024-09-26 - Exact published timestamp: 2024-09-26 - Topics: Risk Management, Brownian Motion, Geometric Brownian Motion - Categories: Risk Management - Source: FE-535 \| Risk Management - Source URL: None The stochastic-process notes move from a random walk to Brownian motion and then to geometric Brownian motion. Brownian increments are normally distributed with variance proportional to elapsed time, while the geometric model keeps a positive asset level by applying the process to log returns. With constant parameters, the geometric Brownian motion solution has the form $$S_t=S_0\exp\left((\mu-\tfrac12\sigma^2)t+\sigma W_t\right).$$ The \(-\tfrac12\sigma^2\) term is the Itô correction. It matters when calibrating a model from observed returns because the mean of the log process is not the same as the mean of the price process. ## How rustc_codegen_clr Connects Rust to .NET - URL: https://sharifhsn.dev/blog/linkedin-2024-09-23-rust-mir-to-dotnet/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2024-09-23-rust-mir-to-dotnet.json - Description: A short guide to Rust's compiler pipeline and a backend that targets .NET. - Date: 2024-09-23 - Exact published timestamp: 2024-09-23T11:47:41.887Z - Topics: Rust, Compilers - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/posts/sharif-haason_rust-panics-under-the-hood-and-implementing-activity-7243949143541391360-ia_9 If you’re a fan of compilers, Rust, or .NET, then Michał Kostrubiec’s [blog post](https://fractalfir.github.io/generated_html/rustc_codegen_clr_v0_2_1.html) about his project `rustc_codegen_clr` is a must-read. This is one of the most exciting and fast-moving projects in the Rust sphere, all driven by the hard work of one talented programmer. Here’s a quick summary of what the project is and why it’s so significant: Rust, like many other modern languages, goes through several steps of compilation where it is compiled into different intermediate representations (IRs) before becoming an executable application. The current translation layer goes like so: Rust -> HIR -> MIR -> LLVM IR -> native code. The key step is between MIR and LLVM IR. This is when the Rust compiler frontend `rustc` shells out to a different compiler backend to perform the actual code generation, or “codegen”. The default is LLVM, the highly popular and versatile Apple-supported backend used by Swift and C/C++ (through `clang`). Using an existing backend allowed the Rust team to focus on frontend development. Instead of compiling efficient code to every possible CPU architecture, they were able to take advantage of LLVM’s existing infrastructure and compile to their IR. Although this works great for the vast majority of use cases, there are some hiccups. LLVM doesn’t support every architecture that the standard C/C++ compiler `gcc` does, limiting Rust’s use in embedded environments. It’s also quite slow because Rust emits a lot of type information by design, which takes a long time to compile since LLVM wasn’t built to optimize that. That’s why efforts have been made to use other backends for codegen, like gcc [2], and there even exists now native Rust backends like Cranelift. [3] `rustc_codegen_clr` [4] does something different. It translates MIR to the C# IR CIL, which allows .NET to compile Rust code. This has two main effects. .NET can use Rust libraries, which are comparable in performance to C libraries while being more robust. Rust can also use .NET libraries, which would enable Rust programs to access much more of the Windows platform, including the Windows App SDK. Rust native modern Windows GUI is just on the horizon! Michał’s blog thoroughly examines the technical details behind his implementation and is worth the read for anyone interested in the project! [1](https://fractalfir.github.io/generated_html/rustc_codegen_clr_v0_2_1.html) · [2](https://lnkd.in/egT4WMJB) · [3](https://lnkd.in/eZaJCvu4) · [4](https://lnkd.in/eZdNpiEr) ## Discrete Distributions and Moments - URL: https://sharifhsn.dev/blog/discrete-distributions-and-moments/ - Structured data: https://sharifhsn.dev/api/posts/discrete-distributions-and-moments.json - Description: For a discrete random variable, the probability mass function gives the mass at each possible value: - Date: 2024-09-23 - Exact published timestamp: 2024-09-23 - Topics: Probability Theory, Discrete Distributions, Moments - Categories: Probability Theory - Source: FE-540 \| Probability Theory - Source URL: None For a discrete random variable, the probability mass function gives the mass at each possible value: $$p_X(x)=\mathbb P(X=x),\qquad \sum_x p_X(x)=1.$$ The CDF is a staircase formed by adding those masses. The notes use a die and the waiting time for a six to show how a geometric distribution arises: if \(X\) is the first roll on which a fair die shows six, then \(\mathbb P(X=n)=(5/6)^{n-1}(1/6)\). Moments are sums in the discrete case. Absolute integrability gives $$\mathbb E[X]=\sum_x x\,p_X(x),$$ and variance measures the squared distance from the mean, \(\operatorname{Var}(X)=\mathbb E[X^2]-\mathbb E[X]^2\). Linearity of expectation is the main calculation shortcut: it does not require independence. ## Martingales and Random Walks - URL: https://sharifhsn.dev/blog/martingales-and-random-walks/ - Structured data: https://sharifhsn.dev/api/posts/martingales-and-random-walks.json - Description: The FE-610 notes describe a martingale as an adapted process whose conditional expected future value equals its current value: - Date: 2024-09-19 - Exact published timestamp: 2024-09-19 - Topics: Stochastic Calculus, Martingales, Random Walks - Categories: Stochastic Calculus - Source: FE-610 \| Stochastic Calculus - Source URL: None The FE-610 notes describe a martingale as an adapted process whose conditional expected future value equals its current value: $$\mathbb E[M_t\mid\mathcal F_s]=M_s,\qquad s\leq t.$$ A symmetric random walk provides the discrete model. Encode heads as \(+1\) and tails as \(-1\); independent increments give zero conditional drift. A Markov process goes further by saying the current state contains enough information about the future, while a martingale says the best conditional forecast is the present value. The notes also introduce first-order and quadratic variation. A smooth function's squared increments vanish in the limit, but a random walk's accumulated squared increments do not. That difference is why Brownian motion needs a new calculus. ## Monte Carlo Risk Simulation - URL: https://sharifhsn.dev/blog/monte-carlo-risk-simulation/ - Structured data: https://sharifhsn.dev/api/posts/monte-carlo-risk-simulation.json - Description: The FE-535 lab notes use Monte Carlo methods to turn a model for returns into a distribution of portfolio outcomes. Simulate the relevant risk factors, revalue the portfolio on eac… - Date: 2024-09-19 - Exact published timestamp: 2024-09-19 - Topics: Risk Management, Monte Carlo, Simulation - Categories: Risk Management - Source: FE-535 \| Risk Management - Source URL: None The FE-535 lab notes use Monte Carlo methods to turn a model for returns into a distribution of portfolio outcomes. Simulate the relevant risk factors, revalue the portfolio on each draw, and summarize the resulting loss distribution. The simulation is only as credible as its inputs. The notes emphasize choosing a horizon, calibrating the distribution, and checking sampling error. More paths reduce Monte Carlo noise, but they do not fix a misspecified dependence structure or a missing stress scenario. For a portfolio with nonlinear instruments, full revaluation is often the cleanest approach. Linear approximations are faster, but their error grows exactly where risk management cares most: large moves and changes in volatility. ## Random Variables and CDFs - URL: https://sharifhsn.dev/blog/random-variables-and-cdfs/ - Structured data: https://sharifhsn.dev/api/posts/random-variables-and-cdfs.json - Description: The FE-540 notes define a random variable as a measurable function from the original sample space to the real line. Measurability means that inverse images of Borel sets are events… - Date: 2024-09-16 - Exact published timestamp: 2024-09-16 - Topics: Probability Theory, Random Variables, CDFs - Categories: Probability Theory - Source: FE-540 \| Probability Theory - Source URL: None The FE-540 notes define a random variable as a measurable function from the original sample space to the real line. Measurability means that inverse images of Borel sets are events in \(\mathcal F\), so probabilities can be assigned to statements such as \(X\leq x\). The distribution of \(X\) is the push-forward probability measure \(\mathbb P_X=\mathbb P\circ X^{-1}\). Its cumulative distribution function is $$F_X(x)=\mathbb P(X\leq x).$$ A CDF is increasing, right-continuous, and tends to 0 and 1 at the two ends of the real line. Interval probabilities follow from differences, \(\mathbb P(x0\), $$\mathbb P(A\mid B)=\frac{\mathbb P(A\cap B)}{\mathbb P(B)}.$$ The total-probability and Bayes formulas update a partition of the sample space when an observation occurs. The notes emphasize that Bayes' rule is a bookkeeping identity for the same joint event, not a new probability model. Independence is the special case in which \(\mathbb P(A\cap B)=\mathbb P(A)\mathbb P(B)\). ## Bessent on Tariffs and the Deficit - URL: https://sharifhsn.dev/blog/linkedin-2024-09-05-bessent-tariffs-and-deficits/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2024-09-05-bessent-tariffs-and-deficits.json - Description: A response to Scott Bessent's arguments about tariffs, inflation, and fiscal policy. - Date: 2024-09-05 - Exact published timestamp: 2024-09-05T11:41:31.361Z - Topics: Tariffs & Trade, Fiscal Policy - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/posts/sharif-haason_bloombergradio-activity-7237424607861878784-wDAo Listened to Scott Bessent, potential Secretary of the Treasury for a future Trump administration, today on Bloomberg Radio talk about his opinions on the fiscal policy of the Trump-Vance campaign and I have some thoughts. He claims that Trump’s tariffs won’t be inflationary because it’s a one-time administrative cost adjustment that will be priced into the market. That might be true in the long-term, but the economic shift of a 10% tariff for the United States would be seismic and might even cause a recession to occur. You can’t compare the effects of specific tariffs like those he instituted on steel to the widespread tariffs he plans to implement. He also deflected from criticisms of Trump’s tax cuts as contributing to the deficit by focusing on the principle of how spending on government programs should be reduced. But this is irrelevant on a fiscal level. On the balance sheet, tax cuts show up the same as government spending. He also claims that tariffs will help pay for it, but as I’ve already noted, there would be massive shifts in the market that would be caused by the tariffs, making the net impact on the deficit hard to predict even if the balance sheet shows a reduction. What do you think about Bessent’s claims? ## Risk Measurement and Portfolio Construction - URL: https://sharifhsn.dev/blog/risk-measurement-and-portfolio-construction/ - Structured data: https://sharifhsn.dev/api/posts/risk-measurement-and-portfolio-construction.json - Description: FE-535 opens by separating the risk-management process from a single risk number: identify exposures, measure them, evaluate the result, and decide how the portfolio should be cons… - Date: 2024-09-05 - Exact published timestamp: 2024-09-05 - Topics: Risk Management, Portfolio Construction, Market Risk - Categories: Risk Management - Source: FE-535 \| Risk Management - Source URL: None FE-535 opens by separating the risk-management process from a single risk number: identify exposures, measure them, evaluate the result, and decide how the portfolio should be constructed. The notes distinguish market, credit, liquidity, operational, and model risk because each fails in a different way. Portfolio construction starts with the joint behavior of returns. Expected return, variance, and covariance determine how a position contributes to the whole portfolio. Diversification reduces idiosyncratic risk only when the exposures are not perfectly aligned; it cannot remove a common market shock. The class frames risk measurement as a decision tool. A statistic is useful only if its assumptions, horizon, liquidity, and tail behavior match the decision it is meant to support. ## Probability Spaces - URL: https://sharifhsn.dev/blog/probability-spaces/ - Structured data: https://sharifhsn.dev/api/posts/probability-spaces.json - Description: The first step in the FE-540 notes is to give random experiments a mathematical home. A probability space is a triple \((\Omega,\mathcal F,\mathbb P)\): \(\Omega\) is the sample sp… - Date: 2024-09-02 - Exact published timestamp: 2024-09-02 - Topics: Probability Theory, Measure Theory, Sigma Algebras - Categories: Probability Theory - Source: FE-540 \| Probability Theory - Source URL: None The first step in the FE-540 notes is to give random experiments a mathematical home. A probability space is a triple \((\Omega,\mathcal F,\mathbb P)\): \(\Omega\) is the sample space of outcomes, \(\mathcal F\) is the collection of events we are allowed to measure, and \(\mathbb P\) assigns their probabilities. The restriction to a sigma-algebra matters when the sample space is infinite. It contains \(\Omega\), is closed under complements, and is closed under countable unions. On the real line, the Borel sigma-algebra is generated by open intervals. A measurable space \((\Omega,\mathcal F)\) is therefore the structure that lets a probability measure talk about events without pretending every subset is measurable. The notes use unions, intersections, complements, and De Morgan's laws repeatedly. They are the set operations that make later statements about random variables and stochastic processes precise. ## What the Price Gouging Prevention Act Says - URL: https://sharifhsn.dev/blog/linkedin-2024-08-27-price-gouging-prevention-act/ - Structured data: https://sharifhsn.dev/api/posts/linkedin-2024-08-27-price-gouging-prevention-act.json - Description: A summary of the 2024 bill's price-gouging language, defenses, and penalties. - Date: 2024-08-27 - Exact published timestamp: 2024-08-27T18:32:34.267Z - Topics: Consumer Protection, Public Policy - Categories: None - Source: LinkedIn - Source URL: https://www.linkedin.com/posts/sharif-haason_monopoly-round-up-price-gouging-vs-price-activity-7234266560796786688-BlTq The newsletter linked here says that Kamala “was not specific on what [price gouging] meant”. There's a bill currently in the Senate called the Price Gouging Prevention Act of 2024, which lays out these details. Is it possible that this bill’s contents are what Kamala is referring to? Here's a quick summary: - The specific language used for price gouging is the following: “It shall be unlawful for a person to sell or offer for sale a good or service at a grossly excessive price, regardless of the person's position in a supply chain or distribution network.” - Small businesses with less than $100,000,000 in gross revenue are exempted from this. - Businesses can make an affirmative defense to the accusation of price gouging by a preponderance of the evidence (i.e. more likely than not) that the increase in price was both not in their control and incurred in the process of procurement, acquisition, distribution, or provision of the good or service. There are also more specific rules in the case of an exceptional market shock, defined as “any change or imminently threatened change in the market for a good or service resulting from [any] cause of an atypical disruption in such market; or any period of time during which the President has declared a major disaster or emergency”. If businesses offer goods at an excessive price compared to what they did before the shock or competing sellers, they cannot have *unfair leverage*, defined as either being large (>$1,000,000 gross revenue/yr), being a critical trading partner, discriminating between equal trading partners, engaging in deceptive practices, or having a dominant position in any market i.e. they have no competition or greater than 30-40 percent of a relevant market. The punishment for price gouging is civil penalties, which for businesses without unfair leverage is a small amount, at most $25,000. For businesses with unfair leverage, the penalty is 5% of their revenues for the previous year. **Bottom line:** The bill targets large businesses that sell goods or services at grossly excessive prices without the affirmative defense that the price increase was not within their control. In addition, during exceptional market shocks, the bill focuses on companies with unfair leverage raising prices relative to other companies, with harsher penalties for those companies as well. Unlike what some critics have said, this is not a price control. It is intended to prevent larger players from using their market share during exceptional market shocks to gouge consumers. The linked discussion: [Monopoly Round-Up: Price Gouging vs. Price Fixing vs. Price Controls](https://www.thebignewsletter.com/p/monopoly-round-up-price-gouging-vs-price-controls). ## Floating Point Architecture - URL: https://sharifhsn.dev/blog/floating-point-architecture/ - Structured data: https://sharifhsn.dev/api/posts/floating-point-architecture.json - Description: Floating point numbers are complicated enough in the abstract. How do systems programmers and hardware designers cope with the complexity of floating point? - Date: 2022-05-31 - Exact published timestamp: 2022-05-31 - Topics: floating-point, Computer Architecture - Categories: Computer Architecture - Source: Archive - Source URL: None Floating point numbers are complicated enough in the abstract. How do systems programmers and hardware designers cope with the complexity of floating point? ## Floating Point Instructions In the general computer architecture, CPUs function by executing a stream of instructions. The set of instructions is defined by the hardware model, such as x86 or ARM. Instructions are typically tiny operations that need to be done extremely often. Some examples of instructions are `mov` to move values in and out of memory, `sub` to subtract numbers in registers, etc. In order to speed up floating point operations, hardware can include floating point instructions instead of letting software handle it. Most of the time, instruction sets only allow floating point operations in the same precision. However, there is a useful optimization that is not often implemented in hardware. A hypothetical instruction to multiply two single precision numbers that results in a double precision number would be very useful. In particular, it would enable for the use of *iterative improvement* algorithms that would otherwise have compounding catastrophic cancellations if the precision was not increased temporarily. ## Subexpression Evaluation The order of operations is very important for preserving invariants in floating point. For example, floating point numbers do not necessarily follow the associative property, so compilers that optimize out parentheses will exhibit undefined behavior. This can also cause issues when numbers with different precisions are operated on. There are two solutions to this problem. One is the solution that languages like OCaml take where all values in an arithmetic expression must have the same type. This means that integers will never be implicitly cast to floating point when they are operated with them, because they are not allowed to be operated with them anyway. However, this can be an overly strict restriction on the types of programs that can be written. Another solution is the one that C takes, which is to establish rules for subexpression evaluation while allowing mixed type expressions. K&R C requires that every operation be done in double precision, but this can cause incompatibility between expressions and a stored value that are in different precisions. A better way to do subexpression evaluation is to set a tentative precision for each expression, then increase the precision as needed in the wider expression. However, this means that the precision of a sub-expression can change with the expression it is embedded in, which is potentially unexpected behavior for a programmer. ## IEEE Conformation Languages like C conform to the IEEE specification, but the ways in which they do it can cause issues in implementation. If the exact rounding operation is implemented in hardware, then all the language has to do is call the appropriate instruction. However, there are some aspects of floating point that are harder to represent. Floating point has a certain *state* associated with it, with rounding mode, flags, trap handlers, etc. This state must be read and written to, and must be preserved across subroutines. `NaN`s also cause issue. The reflexive property is typically taken as invariant since it is in integer math, but it can be disastrous with `NaN`s. One massive consequence of the inclusion of `NaN`s is that floating point numbers *cannot have a total order*. This is because `NaN`s are explicitly unordered with respect to other floating point numbers. This has cascading consequences for the implementation of comparison operators like <, >, and =. ## "Optimization" Compilers make many optimizations to programs in order to increase performance or reduce instruction count. These optimizations are usually made with certain invariants in mind so that the accuracy of the program is not affected. However, because floating point has different invariants than integer math, compilers can make mistakes when optimizing floating point math. For example, ```c float ε = 1; do { ε = 0.5 * ε; } while (ε + 1 > 1); ``` will estimate \\(ε\\). However, a compiler might notice that `ε + 1 > 1` for integer math is equivalent to `ε > 0` and make that optimization, without considering that the expression \\(ε ⊕ 1\\) has a special meaning that is different for floating point numbers than a test for positivity. In general, many algorithms with floating point will exhibit expressions that, upon first blush, seem to be redundant in integer math. Having consideration for the ways in which floating point math can change values is important. ## Exceptions Trap handlers are empowered to be able to access variables in programs. However, computers which have parallel arithmetic may not necessarily be able to easily identify which operation threw an exception. Trap handlers must also be able to identify programs. The reordering of certain instructions can also cause issues. Compilers will often change the order of certain arithmetic instructions when they are exact so that the semantics don't change. However, if the operation traps, then it is not as easy to identify the operation that trapped, since the operation that happened in parallel will modify the trapped arithmetic. One solution to this is *presubstitution* where a user that knows that an exception can create their own handler for an exception and substitute the value themselves beforehand. However, because it goes against IEEE-754, it's unlikely to be proliferated. *All this information comes from the landmark paper [What Every Computer Scientist Should Know About Floating-Point Arithmetic](https://docs.oracle.com/cd/E19957-01/800-7895/800-7895.pdf)* ## IEEE-754 - URL: https://sharifhsn.dev/blog/ieee-754/ - Structured data: https://sharifhsn.dev/api/posts/ieee-754.json - Description: IEEE-754 is the standard for floating point computation around the world. Thus, any real-world discussion of floating point must include it. - Date: 2022-05-31 - Exact published timestamp: 2022-05-31 - Topics: floating-point, Computer Architecture - Categories: Computer Architecture - Source: Archive - Source URL: None **IEEE-754** is the standard for floating point computation around the world. Thus, any real-world discussion of floating point must include it. ## Format IEEE-754 floating point numbers are always in base 2 binary; this makes sense as it minimizes wobble which scales with \\(β\\). It also allows for a special optimization that can add an extra bit of precision without increasing the size of the number. To demonstrate that, we can look at the number \\(001010.0011\\). The normalized floating point representation of this is \\(1.0100011 × 2^3\\). Notice how the digit in front of the decimal point is always 1. It can't be 0, because then we could just shift the decimal point and throw away the zero as we did with the initial number. Therefore, IEEE-754 assumes that the first bit is always 1 and only encodes the bits after the decimal point, granting an extra 1. There is an obvious problem with this: the number 0 cannot be represented. In IEEE-754, if the significand and exponent are all 0, then the number is the special case of 0. IEEE-754 allows for four possible precisions: single, double, single-extended, and double-extended: | Parameter | Single | Single-Extended | Double | Double-Extended | | ----------------- | ------ | --------------- | ------ | --------------- | | \\(p\\) | 24 | 32 | 53 | 64 | | \\(e_{max}\\) | +127 | +1023 | +1023 | +16383 | | \\(e_{min}\\) | -126 | -1022 | -1022 | -16382 | | bits for exponent | 8 | 11 | 11 | 15 | | total # of bits | 32 | 43 | 64 | 79 | The usefulness of extended precision is that it grants a large amount of guard digits for internal calculations, for example in a calculator. There are fast algorithms for common transcendental functions like \\(\log\\), but most of them have a large amount of possible error. By using extended precision internally then rounding to regular precision on display, calculators can quickly compute transcendental functions for users accurately. The exponent value is calculated using a *bias* so it can represent negative values. The initial unbiased exponent is subtracted by \\(2^{e - 1} - 1\\), e.g. 127 for single precision. The reason that the \\(e_max\\) is always bigger than \\(e_min\\) is to prevent overflow when the reciprocal of the smallest numbers are taken. In turn, the reciprocal of the largest numbers will create underflow. In general, underflow is preferable to overflow, for reasons that will be clear when we discuss overflow in more depth. Operations between floating point numbers in IEEE-754 are always exactly rounded. Having a guard digit can reduce error, but it does not guarantee the same result as exact rounding, so it cannot be used. There are methods using two guard digits and a "sticky" bit that can be used for efficient exact rounding, however. Exact rounding is important so that the result of two operations can always be guaranteed to be the same, regardless of hardware. The operations defined under IEEE to be exactly rounded are +, -, ×, /, √, %, and conversion with integers. Conversion between binary and decimal is *not* necessarily exactly rounded, because the most efficient conversion algorithms are not exactly rounded. Transcendental functions are also not specified as exactly rounded because there is no efficient algorithm that works across all hardware. ## `NaN` and ∞ If you've used a calculator before, chances are you have seen both of these terms before. But what do they mean in the context of floating point? Not all bit patterns in IEEE-754 follow the rules that were laid out earlier for significands and exponents. Some aren't even valid floating point numbers! There are special values that are possible: | Exponent | Fraction | Represents | | --------------------------- | -------- | ------------------------------------------ | | \\(e = e_{min} - 1\\) | f = 0 | ±0 | | \\(e = e_{min} - 1\\) | f ≠ 0 | \\(0.f × 2^{e_{min}}\\) (**denormalized**) | | \\(e_{min} ≤ e ≤ e_{max}\\) | — | \\(1.f × 2^e\\) (**normalized**) | | \\(e = e_{max} + 1\\) | f = 0 | ±∞ | | \\(e = e_{max} + 1\\) | f ≠ 0 | `NaN` | The middle row represents normalized numbers, which we are already familiar with. We will discuss each of the other special values. As mentioned earlier, 0 is an exception where the significand and exponent are both 0. But wait, why can 0 be positive *and* negative? IEEE-754 also defines that \\(+0 = -0\\) so the distinction seems useless. The reason both are allowed is because it preserves the sign of infinity when they interact. It also allows for functions that are only defined for positive or negative numbers to be able to safely include 0 in them. *Denormalized numbers* are numbers with 0 as the hidden bit instead of 1. The reason they need to exist and break these rules is because small numbers being subtracted can end up underflowing to 0. It breaks a common invariant in code that the difference between two unequal numbers is not 0, and can cause bugs. Denormalized numbers gradually underflow to 0 so this problem does not occur. It also significantly reduces relative error for very small numbers. The idea of infinity is an important one to measure for floating point. Unlike with integer arithmetic, division by zero is not an immediate error which aborts the operation. Rather, overflow can be an expected result which is handled for by using infinity. If the infinity ends up in a denominator, then the computation can underflow to 0 which can be an expected result. `NaN` stands for "Not a Number" and is perhaps the most unique special value`NaN` does not represent any computable value, and it is "infectious"; any operation with`NaN` will result in an `NaN` and `NaN`s always compare to false, even with other `NaNs`. They are a result of computations like \\(0/0\\) or \\(\sqrt{-1}\\) which are not well-defined in real numbers, and also cannot be represented by infinity. These are all the operations that will produce an `NaN`: | Operation | Production | | --------- | ------------------------------ | | + | \\(∞ + (-∞)\\) | | × | \\(0 × ∞\\) | | / | \\(0/0\\), \\(∞/∞\\) | | % | \\(x % 0\\), \\(∞ \% y\\) | | √ | \\(\sqrt{x}\\) for \\(x < 0\\) | Unlike with denormalized numbers, the value in the significand does not necessarily represent any specific values. Often, some information about the operation will be placed in the significand as a signal. If an operation is done between a real number and an `NaN`, it will not change the significand of the `NaN`. ## Exceptions `NaN` and ∞ allow floating point calculations to continue in the face of special circumstances without aborting immediately. However, this behavior is not always appropriate. Implementations of IEEE-754 typically have **trap handlers** that will handle **exceptions** generated by such circumstances. There are five classes of exception that can each be set by a status flag: overflow, underflow, division by zero, invalid operations, and inexact operations. Invalid operations are any that involve `NaN` except for those when one of the operands are already `NaN`. Operations in IEEE-754 must be exact, so if an operation is performed that is inexact, it will raise an exception. However, this raises an issue. Inexact exceptions can happen extremely often, and telling the OS to summon the trap handler every time harms performance. In order to prevent this, a software flag is typically enabled upon inexact exception to mask off future exceptions until the flag is reset. Although trap handlers can abort the algorithm, they can also exhibit other behavior which can aid algorithms in being more efficient and accurate. For example, IEEE-754 will wrap around overflowing numbers by dividing the computed result by \\(2^α\\). \\(α\\) is 192 for single precision and 1536 for double precision. This is useful in partial products when the operation might overflow and underflow at several points. By allowing the products to continue without aborting, the final product might cancel out under/overflows and be in range without having to stop the OS at every point. ## Rounding We stated earlier that round to even is the best and most commonly accepted method of rounding. This is the default mode of IEEE-754. However, there are other acceptable rounding modes that can be enabled for certain operations. These are round to 0, to +∞, and to -∞. These modes turn rounding into floor/ceiling operations, which are useful for computing intervals. We can represent floating point results as an interval between two numbers where the first is rounded to -∞ and the second is rounded to +∞. The exact result is somewhere between those two numbers. This representation can be useful to get an idea of how wide your error is when calculating a result. If the result with single precision has a very wide interval, then you might redo the result in double precision and so on to shrink the interval to some acceptable level of error. *All this information comes from the landmark paper [What Every Computer Scientist Should Know About Floating-Point Arithmetic](https://docs.oracle.com/cd/E19957-01/800-7895/800-7895.pdf)* ## Floating Point - URL: https://sharifhsn.dev/blog/floating-point/ - Structured data: https://sharifhsn.dev/api/posts/floating-point.json - Description: Floating point representation is a complex topic full of nuance. To understand it, we must focus on the abstract method used to represent non-integer numbers. - Date: 2022-05-28 - Exact published timestamp: 2022-05-28 - Topics: floating point, Computer Architecture - Categories: Computer Architecture - Source: Archive - Source URL: None **Floating point representation** is a complex topic full of nuance. To understand it, we must focus on the abstract method used to represent non-integer numbers. ## Numbers in Computers Math is infinite. Or at least it can be. Numbers like \\(5\\) can be represented finitely, but a number like \\(π\\) is not so simple to represent. \\(3.14159265358979\dots\\) but that still isn't enough. It'll never be enough. Unfortunately, computers are not infinite. Although we would like to hold every digit of \\(π\\) in our computer's memory, the fact is that there will always be limits to how numbers can be represented in computers. Integers are numbers like \\(3\\), \\(15\\), \\(-154\\), etc. They have no decimal point and can be negative. Representing these numbers in a computer is fairly simple, with the limit only being placed on the size of the number. However, not all real numbers play so nicely. How do we represent the value \\(12.5\\) in memory? We can't have half bits. Can we go infinitely precise on decimals, like \\(7.2028301324\\)? What about the size? How do we set these bounds? ## Floating Point We have to have a different standard to define how these numbers will be represented in a computer. This standard is called **floating point**. Generally, a floating point representation is composed of three parts: the **base** \\(β\\), the **significand** \\(p\\), and the largest/smallest allowed exponents \\(e_{max}\\) and \\(e_{min}\\). This is the complete formula for floating point (ignoring specific details): $$ \lceil \log_2{(e_{max} - e_{min} + 1)} \rceil + \lceil \log_2{β^p \rceil + 1} $$ For example, let's say we had a floating point representation in base 2 which allowed for exponents of sizes up to \\(+128\\) and down to \\(-127\\), and had twenty-three bits of precision for the significand. This representation would need 32 bits. A number written in floating point representation might look like this: $$ 1.010001 × 2^{7} $$ where the first part is the significand, the base of the exponent is the base, and it's raised to its exponent. Floating point numbers are written like scientific notation to be *normalized*, where there is only one nonzero digit in front of the decimal point—this will become important for IEEE 754. ## Imprecision I mentioned bits of precision, which is finite. Some decimal numbers can't be represented by a finite amount of precision. For example, the number \\(0.3\\) is finitely representable in base 10, but becomes \\(\overline{1.001} × 2^{-2}\\) in binary. This number will *never* be perfectly represented in binary, so we have to have approximations. In order to understand how much error we encounter, we must be able to measure it. There are two techniques to measure error in floating point: `ulps` or "units in the last place" and *relative error*. Let's use an example here, with \\(β = 10\\) and \\(p = 3\\). We want to approximate \\(2.781828\\) to a floating point number with this precision. Since we can only encode 3 digits of precision, we end up with \\(2.78 × 10^0\\). If we imagine a special decimal point after the precision stops for floating point, then the difference between these two is \\(0.1828\\). This is difference for `ulps`. **The closest floating point number can still have a `ulps` error of up to \\(\frac{β}{2} \cdot β^{-p}\\). This number is known as the machine epsilon \\(ε\\).** Relative error is a familiar concept in most sciences. It is simply the difference between measured and actual, divided by the actual value. In this case, our floating point approximation \\(2.78 × 10^0\\) is the "measured" value and our real number \\(2.781828\\) is our actual value. In this case, it is \\(0.001828 / 2.781828 = 0.0006\\). Since this number can be small, it is often expressed in terms of \\(ε\\). In this example, \\(ε = 5 × 10^{-3} = 0.005\\) so our relative error is \\(0.13ε\\). Since we know the maximum `ulp` error for the closest floating point number, we should find the same for relative error. One difference about relative error is that it changes based on the size of the numbers, not just the absolute error in the significand. The largest possible error is \\(\frac{β}{2}β^{-p} × β^e\\). However, the relative error changes based on the size of the real number in question, and all real numbers within the same exponent range will have this error. So, the relative error for a number closer to \\(1.0 × β^e\\) will be larger than for a number closer to \\(β × β^e\\). This variation is called **wobble**. The relative error is always bounded by \\(ε\\), as in the prior example. Depending on the significand, however, the wobble can be up to \\(β\\) for the same exponent. Crucially, this wobble can be expressed in either relative error or `ulps`, as long as the other is held fixed. Typically, relative error is used, because it is more useful for compounding operations whereas `ulps` can vary wildly within \\(β\\). An important concept to understand here is **contaminated digits**. These are the least significant digits of a floating point number which may be error-prone and are therefore untrustworthy. The number of contaminated digits is \\(\log_β{n}\\) where \\(n\\) is the factor of the relative error to \\(ε\\), like \\(0.13\\) in the earlier example. ## Error Mitigation We have found how to measure error; now how do we reduce it? One method is to use **guard digits**. When performing calculations, a computer can "extend" the calculation by a few digits so that the extended digits are contaminated by the floating point operation. These digits will then be rounded out of the final result, reducing the contamination of the result. Without a guard digit, the relative error can be large as \\(β - 1\\), which can put every digit in error! However, adding just *one* guard digit bounds the relative error to *less than \\(2ε\\)!* Another method is to reduce the number of **cancellations**. A cancellation occurs when nearby quantities are subtracted. When this happens, the most significant and uncontaminated digits cancel out and the less significant, more contaminated digits are left. Cancellations can seriously magnify rounding errors in a series of cancellations; this is called *catastrophic cancellation*. However, cancellation can also be *benign*. If the two quantities being subtracted are exactly known, then the subtraction will have a tiny relative error if done with a guard digit. Sometimes, we can rearrange formulas to have less or more benign cancellations. For example, the formula \\(x^2 - y^2\\) has a catastrophic cancellation because \\(x^2\\) and \\(y^2\\) both suffer from rounding error. However, this formula can be rearranged to \\((x ⊕ y) ⊗ (x ⊖ y)\\). The cancellation is now benign because \\(x\\) and \\(y\\) are presumably exact values without rounding error yet. > Notice that the operands in that formula are circled. This is a notation to indicate that these operations are performed by a computer and therefore may accrue rounding error, whereas the ordinary operands are used for exact calculations. However, the impact of making a cancellation benign can be limited if the inputs to the equation are already inexact. This is common when converting between decimal and binary numbers. **It is always worth eliminating a cancellation, though.** ## Exact Rounding Although guard digits mitigate error, they will in many cases give a different value than the **exactly rounded** result. That is the result if the floating point number was computed exactly, then rounded to precision. Many algorithms require exact rounding to work properly. The nature of rounding is controversial. The most commonly accepted rounding mode is *round to even*. This will round numbers that end in the half digit to whichever direction makes the number even. For example, \\(12.5\\) will round to \\(12\\), not \\(13\\) because \\(12\\) is even. This achieves the result of rounding up and down being equal chance for the half digit, 5 in this case. Exact rounding is useful for holding certain invariants in floating point calculation. For example, one way to increase precision in calculations is to split a multiple precision number into an array of single precision numbers. Adding these numbers back together with exact rounding will recover the multiple precision numbers. We can represent a multiplication of two double precision floating point numbers \\(x\\) and \\(y\\) using this method. We can split \\(x\\) into \\(x_h\\) and \\(x_l\\), each of which are single precision, and \\(y\\) into \\(y_h\\) and \\(y_l\\) likewise. Then, we can represent the multiplication as so: $$ x \cdot y = (x_h + x_l)(y_h + y_l) = x_hy_h + x_hy_l + x_ly_h + x_ly_l $$ In this way, a multiplication of two double precision numbers can be represented as the sum of multiplications of single precision numbers. Because the numbers are exactly rounded, the number is completely recoverable. This means that double precision multiplication can be possible on a computer which only supports single precision multiplication. *All this information comes from the landmark paper [What Every Computer Scientist Should Know About Floating-Point Arithmetic](https://docs.oracle.com/cd/E19957-01/800-7895/800-7895.pdf)* ## Network Security - URL: https://sharifhsn.dev/blog/network-security/ - Structured data: https://sharifhsn.dev/api/posts/network-security.json - Description: The internet has taken over the world, but with it have come malicious actors. People want to spy on you, scam you, steal from you, or harm you through the internet. Through networ… - Date: 2022-04-22 - Exact published timestamp: 2022-04-22 - Topics: Internet Technology - Categories: Internet Technology - Source: Archive - Source URL: None The internet has taken over the world, but with it have come malicious actors. People want to spy on you, scam you, steal from you, or harm you through the internet. Through **network security**, we can shield potential victims from at least some of these harms. ## Attacks Attacks on a network can be performed *passively*, for example in eavesdropping. The network cannot tell the difference between this passive attack and ordinary traffic because no more traffic is being injected. *Active* attacks, by contrast, are done by changing the kind or amount of traffic. This typically manifests as a **DoS** **(Denial of Service)** attack. This is usually a significant increase in traffic which shuts down the network due to congestion. If only one host is attacking the network, then the network can easily shut a particular malicious host out. DoS attacks are usually distributed, hence **DDoS** **(Distributed Denial of Service)**. Teardrop attacks are a modern form of DDoS. Instead of sending lots of useless traffic, all they do is open a new connection with a socket. Even though this is a small operation, when done on large scale, it can tax the memory of a server. We mentioned previously that in order to compensate for fragmentation, routers will keep a memory buffer to store fragments. This can be exploited by malicious packets that are all fragments which will run out the memory of the router. The three main aspects of network security are **authentication**, **confidentiality**, and **integrity**. - Authentication: How do you prove that someone is who they say they are? - Confidentiality: How can you hide data from prying eyes? - Integrity: How can you prove that data has not been altered midstream? In order to protect these, we need a **security model**. 1. An algorithm to ensure security that must be well-known. 2. Trusted parties know how to generate secret information. 3. Distribute that information through a secure method. 4. Specify a protocol for trusted parties to decrypt data. ## Encryption **Encryption** or **cryptography** is the encoding of a message in such a way that only the communicating parties can interpret it. That way, if a malicious party tries to read the message, they will be unable to understand it because they cannot decrypt it. Encryption has two sides, the encryption algorithm and the decryption key. The algorithm should be publicly available in order to prove that it works. When you feed the decryption key into the algorithm, it will spit out a string associated with only that key. The secret key itself should be long enough that the algorithm cannot easily be broken, but also short enough so that it can be transmitted easily. Fun fact: before minimum password lengths were instituted, the most common password was "god". The most basic kind of encryption algorithm is a substitution cipher. By simply replacing letters in a message with another letter with a known mapping, a message can be encoded and decoded. However, it is relatively trivial to break this cipher because the English language has patterns that can identify certain letters better than others. You can increase the complexity of this algorithm through polyalphabetic encryption, where the cipher itself changes after every letter. Famously, this is how the Enigma cipher used by the Nazis during World War II worked. Transposition ciphers will change the *order* of the words as well. This might seem like a really good cipher, unfortunately it can similarly be broken by looking at the structure of language. In order to actually encrypt data, real-world encryption algorithms use a combination of these techniques. But how do we generate the decryption key? There needs to be one for every user, but they need to be as random as possible and nigh impossible to crack, so we can't use plaintext passwords. We can use *pseudorandom number generators* in order to create a key. A pseudorandom number generator is an algorithm which, starting with some seed, will produce a string of numbers based on that seed which *appear* to be random. The seed will typically based on some environmental noise that can't be easily predicted, such as the time of day at which the algorithm runs. The generator should be *highly* sensitive to changes in seed. ## Symmetric Key Using the same key, two parties should be able to both encrypt and decrypt the message. By applying the key to the message, the message should be encrypted, then decrypted. There are two main kinds of symmetric ciphers: *stream ciphers*, and *block ciphers*. As the name implies, stream ciphers do bitwise operations on the message and the key in order to get the ciphertext. The bitwise operations can be symmetric, like ⊗. One example of a stream cipher is RC4 which can be used in SSL. Block ciphers have a 1-to-1 mapping of certain blocks of \\(k\\) bits to other blocks, typically 64 bits. This is a good method, but the problem is that the number of mappings balloons massively: in general, it is \\(2^k!\\), which is absolutely massive! It is both exponential *and* factorial. To simulate such a table, you can use a pseudorandom function. ## DES The **Data Encryption Standard** is the US encryption standard. It uses a 56-bit symmetric key for 64-bit plaintext input, and it is a block cipher. It is fairly secure for small attacks, but it can be brute force decrypted in less than a day. Its advantage is that there is no analytic exploit in the standard. In order to make DES more secure, we can increase the number of keys to 3, making **3DES**. The message is encrypted, decrypted, then encrypted again when run through all three keys. Effectively, there are \\(56 × 3 = 168\\) bits to crack, which is much stronger than 56. ## Replay Attacks This method of symmetric key is vulnerable to a certain kind of attack: the **replay attack**. Let's say I encrypt a message and mail it to my boss. My boss decrypts the message, and sends me back a confidential message, so we can pass information to each other. But my boss doesn't know my address, he only knows that he can trust me because the message is encrypted using our key. What if a malicious hacker somehow obtains my encrypted message? If she sends that message to my boss, then my boss will email the hacker back with the confidential information because he thinks that only I have the ability to encrypt messages in that way. The way to protect against this is by using a **NONCE** value, which is a **N**umber used only **ONCE**. Let's say I write a random number on my message every time I send a message to my boss. The hacker obtains my message and sends the same message later, with the same random number. My boss knows that the NONCE value cannot be the same and therefore does not trust the hacker with our confidential information. The NONCE doesn't have to be a random number. Many times, it can be a challenge query like asking for my mother's maiden name. Having multiple challenges that are cycled through adds an extra layer of protection. ## RSA Nowadays, almost all cryptography is done by having two keys: a **public key**, and a **private key**. The public key is used to *encrypt* messages, and private key is used to *decrypt* messages. The public key is, as might be obvious, publicly available. Someone who wants to send a secure message to me can use my public key to encrypt it. However, only I can use my own private key to decrypt it. Importantly, *you should not* be able to determine the private key by analyzing the public key. They should be wholly different in nature. **RSA** is the most commonly used implementation of this kind of cryptography. To encrypt a message, take the message and raise it to the power of the public key, `mod` some \\(n\\). To decrypt said bit pattern, raise it to the power of the private key, again `mod` the same \\(n\\). Authentication can be done using the passing of challenge messages. A new secret key will be created which is encrypted and decrypted in order to encrypt the data. One problem with RSA is that it's really expensive and slow computation-wise, which is good for security but bad for convenience. In practice, most uses of RSA encryption only use the actual algorithm to encrypt some secret key, which is then used as a simpler cipher for the data between the two parties. ## Hash **Hashing** is the process of creating a unique-ish value from some data that can be used to verify that data as a signature. It should be fast, not easily irreversible, and not have collisions i.e. no two hashes should be the same. The two most common hash functions are **MD5** and **SHA-1**, which are 128 bits and 160 bits, respectively. ## A Crash Course in Python - URL: https://sharifhsn.dev/blog/python-crash-course/ - Structured data: https://sharifhsn.dev/api/posts/python-crash-course.json - Description: Python is a powerful jack-of-all-trades language that is much less limited than OCaml. - Date: 2022-04-14 - Exact published timestamp: 2022-04-14 - Topics: Principles of Programming Languages - Categories: Principles of Programming Languages - Source: Archive - Source URL: None Python is a powerful jack-of-all-trades language that is much less limited than OCaml. ## Expressions Arithmetic works in a very simple and clear way in Python, basically like a calculator. ```python 2 # 2 2 + 4 # 6 (2 + 5) * (3 - 5) # -20 ``` Unlike in OCaml, operations can be performed between floats and integers with no issues; Python will simply change the types behind the scenes. This is called **type coercion**. ```python 2 + 3.5 # 5.5 ``` This causes an interesting side effect: what is the return type of this function? ```python def add(a, b): return a + b ``` In OCaml, the equivalent function would have type `int -> int -> int`, since the operation `+` is only valid for `int`. However, this function is actually polymorphic in Python! In OCaml, you would call it type `'a`. In Python, this type is called **Any**, and operation changes based on what the types are. The reason that this is possible is because *everything* in Python is an object. Essentially, the variables `a` and `b` are just boxes that could contain anything, and Python only checks whether the operation `+` is defined for the two variables at runtime. ## Strings String manipulation is a very common operation in Python, so there are some very useful ways to handle strings built into the language. ```python "hello " + "world" # "hello world" "hello" * 3 # "hellohellohello" ``` You can also convert types to strings very easily. ```python str(5) # "5" str(3.5) # "3.5" ``` And vice versa. ```python int("5") # 5 float("3.5") # 3.5 ``` These are special built-in functions to make our lives earlier. ## Variables Like with OCaml, there is no need to specify types as in Java. Unlike OCaml, Python does not use type inference. Instead, every variable can contain any type in a box as mentioned earlier. ```python a = 3 b = "hello" ``` Variables in Python work differently under the hood than other languages. A variable is essentially just a name for an element. ```python c = b ``` `c` here is not just equal to `b`... it is actually `b` itself! And any change you make to `b` will reflect in `c` because of that. Actual copies must be explicit. ## Slices **Slices** are one of the most powerful tools in Python. In fact, it might be what Python is most known for and most useful for. Slicing allows for powerful manipulation of list-like data. Ordinary list access uses bracket notation with a single number to access a single element. Slice notation works similarly, but it allows returning multiple elements as a sub-list of the original list. ```python x = "hello world" x[1:7] # "ello w" ``` You might notice that I just performed this operation on a string; didn't I just say that slices work on list-like types? In Python, strings are just fancy lists of chars!. ## Tuples and Lists A tuple is an immutable set of multiple elements, just like in OCaml. Slicing operations work the same way in that they return a subsequence tuple. There's a weird side effect where if you slice in a way that returns a single element, you can get a single-element tuple. Lists in Python are much more flexible in OCaml, as they can be heterogenous with any type within. They are also mutable, so elements of lists can be reassigned. Slicing can superpower this assignment by reassigning multiple values at once ## Control The traditional `if` statements are back, but with a bit of a twist. Python has *significant whitespace*, which means that the amount that you indent by affects the actual execution of the code. This is quite rare. ```python if x == 15: y = 0 # this tab is mandatory! x = 3 # this line is outside the if block ``` There is a special keyword called `pass` which exists to allow empty blocks. Normally, this construct is forbidden: ```python if x == 15: x = 3 # error! ``` But by using `pass`, this is possible: ```python if x == 15: pass # does nothing x = 3 ``` Like in C, boolean evaluation is 0 for false, true for everything else. However, you *cannot* assign a variable in a conditional! So this C construct would not be allowed: ```c while ((int x = some_func()) == 0) { // this is allowed in C } ``` ```python while (x = some_func()) == 0: pass # this is not allowed in Python! ``` ## Functions Functions are defined using the `def` keyword. ```python def fac(n): if n < 0: return "negative!" elif n == 0: return 1 else: return n * fac(n - 1) ``` Functional programming can be used similarly to OCaml where everything is a function. ```python def compose(f, g): def foo(x): return f(g(x)) return foo ``` ```ocaml let compose f g = let foo x = f g x in foo x ``` ## Disks - URL: https://sharifhsn.dev/blog/disks/ - Structured data: https://sharifhsn.dev/api/posts/disks.json - Description: Persistent storage is essential to every system that we can think of, and we can hardly think of a computer without it. However, managing disks is not as simple as it might appear … - Date: 2022-04-14 - Exact published timestamp: 2022-04-14 - Topics: Operating Systems Design - Categories: Operating Systems Design - Source: Archive - Source URL: None Persistent storage is essential to every system that we can think of, and we can hardly think of a computer without it. However, managing **disks** is not as simple as it might appear on a surface level. ## I/O Before we dive deep into disks, we must first consider the general concept of **input/output**, or **I/O**. This is a general class of devices which externally connect to the computer and provide input and output. They include keyboards, mice, displays, speakers, etc. and are essential to any typical use of a computer. All of these devices are very different, so how does the operating system communicate with each of them? In order to talk to external devices, the OS treats every device as a **canonical device** with three components: a status register, a command buffer, and a data buffer. The status register tells the OS what the current status of the device is. There might be many types of status for a device, but the basic one we will consider is `BUSY`. If a device is busy, the OS will not command it to do anything and will wait until it is free. Once it's free, the OS will put a command in the command buffer, and optionally some relevant data in the data buffer. The device will become busy again performing the command that the OS has given it. Note that the device here is agnostic to the actual use of the device. The OS uses the same general procedure for every device. One nice thing about I/O theoretically is that because its duties are separate from the computer, it can perform operations while the CPU is busy doing something else, increasing performance. However, this might not necessarily be the case. In order to initiate I/O, the CPU must, as previously stated, modify the command and data buffers. This can use up valuable CPU time. In order to mitigate this, most computers have a separate **DMA engine** for direct memory access. This engine will communicate with I/O while the CPU is busy doing something else. Although we do have this canonical device, the computer still needs to know what the different commands are and the type of data to put into the hardware. In order to communicate with the hardware effectively, all modern operating systems use **device drivers**. These are essentially low-level wrappers to the device which provide a convenient API for operating systems to use. In fact, almost all code in an operating system is devoted to these device drivers! ## Hard Disks For most of computing history, **HDDs** or hard disk drives have been the hardware for persistent storage, and their internal structure is important to making sure they are performant. Recently, **SSDs** or solid state drives have increased in popularity for faster storage access. However, they are still expensive for large amounts of data and HDDs are still used commonly, so let's discuss them. The basic hardware for a hard disk is a platter with several rotating wheels and an arm which moves between each wheel. Most hard disks are composed of thousands of such platters. Each wheel is sectioned into blocks of data. In order to access the data, the arm must **seek** the correct wheel and the wheel must **rotate** to bring the correct data block for the arm to read, which will **transfer** the data to the computer. These three steps form the basis of reading and writing from a hard disk. Seeking and rotating are slow and tend to be the bottleneck of access, while transferring is typically fast. The time taken for I/O is defined as $$ T_{I/O} = T_{seek} + T_{rotation} + T_{transfer} $$ and the rate of I/O is $$ R_{I/O} = \frac{Size_{Transfer}}{T_{I/O}} $$ These facts mean that when composing data, we want to reduce the amount of seeks and rotations as much as possible. This makes **sequential** data very important. When data is accessed sequentially, it tends to be on the same wheel and next to each other, so no seeks and very little rotation needs to be done. ## Scheduling Like with CPU cycles, disk requests can also be scheduled and reordered. This becomes important when trying to optimize for sequential data instead of random access. The most obvious way to schedule a disk is to schedule requests that have the shortest seek time from the current wheel, so either on the same wheel or a close-by one. This increases performance, but it will starve requests that happen to be far away. We want to have some amount of fairness in our scheduling. The currently accepted way to do this is **SCAN** or **C-SCAN**, also known as *the elevator method*. Essentially, the disk will be swept from beginning to end to check for any requests. If there is a request, then it will be performed. This preserves sequential data as the sweep goes in order, but makes sure that far away requests are also done. C-SCAN is an improvement over SCAN. Where SCAN simply goes back and forth like a real elevator, C-SCAN will treat the disk like a circle and sweep in the same direction. This prevents the sectors at the ends from being starved. ## Particles - URL: https://sharifhsn.dev/blog/particles/ - Structured data: https://sharifhsn.dev/api/posts/particles.json - Description: Tiny particles are affected in different ways than we are by waves. There are several different effects that must be discussed. - Date: 2022-04-12 - Exact published timestamp: 2022-04-12 - Topics: Physics - Categories: Physics - Source: Archive - Source URL: None Tiny particles are affected in different ways than we are by waves. There are several different effects that must be discussed. ## The Photoelectric Effect Many people believe that Albert Einstein won the Nobel Prize for his discovery of special relativity, as that is what he is most well-known for. However, his actual claim to fame in that respect had to with the **photoelectric effect**. He found that when light of a certain wavelength is incident on certain metals, electrons are ejected from the surface of the metal. The number of photoelectrons emitted, the photocurrent, increases with intensity, though the kinetic energy does not. The kinetic energy only increases with an increase in energy of the light i.e. a decrease in the wavelength. Above a certain wavelength, no electrons are ejected. $$ \frac{hc}{λ_0} = W_0 + KE_{max} $$ where \\(W_0\\) is the work function for the metal in question. You can think of it as an intrinsic quality of the metal. \\(λ_0\\) is the cutoff wavelength after which no more electrons will be ejected. ## Pair Production/Annihilation When a high energy photon collides with a nucleus, it may "split" into an electron and a positron. When this split happens, the charge, momentum, and the mass-energy of all parties must be conserved: $$ E_{photon} = KE_{positron} + KE_{electron} + 2E_0 $$ where \\(E_0\\) is the rest mass-energy for each particle given by \\(E = mc^2\\), doubled since the electron and positron are identical. The opposite operation can occur when an electron and a positron combining will annihilate the pair and create two identical photons with opposite momentum. The reason that there must be two photons and not just one is that the momentum must equal 0, so there has to be opposite momentum. ## de Broglie Wavelength All particles exhibit wave-like characteristics, which can be demonstrated by the double-slit experiment. The wavelength of a particle in this instance can be determined from the following equation: $$ λ = \frac{h}{p} = \frac{h}{mv} = \frac{h}{\sqrt{2mKE}} $$ The wavelength of a photon is a little different, since photons are completely massless. The momentum and energy of a photon are related by a simpler equation \\(E = pc\\), which applies to all massless particles, not just photons. ## Uncertainty One key problem for physicists studying particles is the **Heisenberg Uncertainty Principle**. In layman's terms, the observation of a particle fundamentally changes its quality, so it can never be truly measured. In mathematical terms, the accuracy of a measurement of a position's momentum and position at the same time is limited: $$ (Δp_y)(Δy) ≥ \frac{h}{4π} $$ and the same function exists for simultaneous measurements of energy and time. ## Multimedia Networking - URL: https://sharifhsn.dev/blog/multimedia/ - Structured data: https://sharifhsn.dev/api/posts/multimedia.json - Description: Streaming multimedia is a relatively new concept in internet technology, but now has become central to the daily media consumption of people around the world. - Date: 2022-04-07 - Exact published timestamp: 2022-04-07 - Topics: Internet Technology - Categories: Internet Technology - Source: Archive - Source URL: None Streaming multimedia is a relatively new concept in internet technology, but now has become central to the daily media consumption of people around the world. ## Multimedia In contrast to the ordinary messages that encode text, **multimedia** streaming uses audio or video. Most of the traffic on the internet today is dominated by multimedia streaming, so it's a big deal. Multimedia streaming is different from other kinds of streaming in that latency is of greatest importance, but it is also loss tolerant; if the streaming video is a bit distorted, it's not a big deal. When the delay hits above 150 ms users begin to complain. VoIP is a special case where there is media being streamed *both* ways, where two people speak on the same connection through data networks. We must handle signals differently because we don't have a simple Huffman coding like with text. There are three elements at play that determine the quality of the signal: - Sampling: how often do you sample the signal for data? - Quantization: how many levels/bits represent each sample? - Compression: how much are the quantized values compressed? In order to minimize the amount of data being sent, video streaming will often only send the changes from one frame to the next. Otherwise, the level of fidelity would be infeasible on current internet technology. ## Audio Audio is generally ranges from 20 Hz to 22.05 kHz, and speech goes between 200 Hz and 8kHz. Sampling generally occurs up to 8kHz to capture speech, the most important element of audio in general, though it is possible to capture 16kHz for high-fidelity audio. Quantization can either be 8 bits or 16 bits, which means there are either \\(2^8\\) or \\(2^{16}\\) levels. That's a difference between 256 and 65,536! The bit rate of samples typically depends on the type of audio being sent. For simple VoIP it will typically be 64Kbps, because it quantized at 8 bits and sampled 8k times per second. Audio will generally be 705.6Kbps from 16 bits of quantization and 44.1k samples per second. Stereo is the most fidelity because it is audio in both ears, so double the bit rate at 1.4112 Mbps. By removing frequencies that the human ear cannot distinguish, the amount of data can be reduced dramatically without altering the end sound. ## Video Images are 2D arrays of pixels at heart. The amount of pixels is *resolution*, which is typically expressed through a 16:9 aspect ratio as with 1080p or 4K/UHD. Every color has 8 bits, and there are three colors RGB so each pixel will contain 24 bits. That means the raw data for a color image even of size 320×240 will be 7680 bytes! That's a lot for one image. Compression through JPEG, Gif, etc. is typically used to help here. Video is composed of images that are displayed at a certain *frame rate*. Movies run at 24 FPS, though some rare movies display at 48 FPS. Video games are commonly at either 30 FPS or 60 FPS, but can go even higher. If we imagine a 4k television displaying color images at 60 FPS video, the amount of data transmitted would be \\(4096 \cdot 2160 \cdot 60 \cdot 24\\) which is 12 Gbps!!! We obviously need some kind of compression. ## Streaming The client will typically downloading the first portion of the multimedia and begin consuming it. The rest is downloaded while the client consumes the data, so you don't need to wait for the whole file to be downloaded. If you've ever watched a YouTube video, you can see this in real time as the gray bar overlaid on top of the timeline of the video. The client has a buffer to store that downloaded video which is constantly drained to show the video and filled at times corresponding to network delays. However, this can cause *buffering* when the network delay is variable and the client consumes all of its data before the server can produce more. The video will be buffered until more data is received, and you see a still frame on the video while a loading symbol plays over it. ## RTP UDP streaming is the traditional type of streaming. As multimedia does not necessarily need to be reliable, UDP is better for reducing latency which is most important for the consumer. The packet has the timestamp of the multimedia so it can be ordered correctly in the client buffer. It also needs certain information like the encoding or whether the audio is for left or right microphone. The protocol used for multimedia that incorporates these is called **RTP** or **Real-time Transport Protocol**. RTP is a flexible protocol that allows for many types of multimedia to be transmitted. The header includes a timestamp field which increments per sample, which allows for correct timing. It also uses sequence numbers to distinguish between packet loss and silence. ## HTTP Most people stream video over their browsers nowadays, which uses HTTP. You might remember that HTTP uses TCP, which is slower than UDP but more reliable. The streaming is slower but more convenient. ## DASH **Dynamic Adaptive Streaming** allows for good bandwidth over HTTP. The client is the one that adjusts the request rate from multiple content formats/encoding. For example, the client can request data at a lower bit rate or quality like 480p if it feels that 720p is too slow. The MPD or Media Presentation Descriptor is sent over HTTP and gives information about the different format segments. The client can then request each segment, and the segments can come from different ASes, which works well for CDNs. This is how ads are served over YouTube; the client is the one that decides whether to get an ad or not. This is also why blocking ads on YouTube can be done by an extension to your browser. ## DNS The ID for every YouTube video is a BASE64 11 character string, giving a space of \\(64^{11}\\) possible videos. The ID is mapped to 192 DNS host names so that multiple servers can be accessed for videos. This ID is a special hash of the video ID. Multiple ASes can broadcast the fact that they are all reachable to a BGP router, and a BGP router can then choose the shortest path. This is called **anycast**. This means that for accessing video servers, you can pick the country where the video is hosted quite easily. Let's say you want to access a Japanese video in Japan. It's obviously best to access the Japan server, so that's the closest. However, all 192 servers that host the video will broadcast, so that if the Japan server goes down in an earthquake or something, the other servers are still available for the BGP router to pick from, so they can access, for example, the South Korean server. ## VoIP Traditional telephones work over PSTN (Public Switched Telephone Network), which is autonomous and only designed for telephones. However, as the Internet has gradually taken over the world, a demand was created to include telephony in the Internet. Thus, the **VoIP** (Voice IP) protocol was born. The voice sound waves are converted into packets and transported as audio through lossy RTP, similar to other forms of multimedia. In order to establish the connection to begin with, **SIP** or Session Initiation Protocol sets up the VoIP session between two hosts. SIP URLS are similar to email addresses that identify users on a network. Most of the work is lent to the caller in order to type in some identifying information as they typically have a display/keyboard. The URL is mapped to an IP address similar to DNS and email. ## Control in Prolog - URL: https://sharifhsn.dev/blog/prolog-control/ - Structured data: https://sharifhsn.dev/api/posts/prolog-control.json - Description: Conditional statements are very powerful and are in used in almost every language. How does Prolog implement the same idea? - Date: 2022-04-05 - Exact published timestamp: 2022-04-05 - Topics: Internet Technology, Principles of Programming Languages - Categories: Principles of Programming Languages - Source: Archive - Source URL: None Conditional statements are very powerful and are in used in almost every language. How does Prolog implement the same idea? ## Logic + Control The basic idea behind a control statement is that there is *logic*, such as facts, rules, and queries, which are composed of clauses/**goals**, as well as *control*, which is how Prolog chooses the logic among several options. The control is implemented by the order of the facts/goals, which is **extremely important**. Prolog will *always* choose the first applicable rule in a program, and it will *always* choose the leftmost clause/goal in a query. Brevity is key. If there are two rules that do the same thing, or one rule that is implied by another rule, you should probably take it out unless executing something more than once is your explicit intention. ## Abstract Interpreter **A substitution \\(σ\\) is a finite set of pairs of terms \\(\\{X_1/t_1, ..., X_n/t_n\\}\\) where each \\(t_i\\) is a term and each \\(X_i\\) is a variable such that \\(X_i ≠ t_i\\) and \\(X_i ≠ X_j\\) if \\(i ≠ j\\).** An empty substitution is denoted by the letter \\(ε\\). Some important rules: - A variable cannot substitute itself e.g. \\(Z/Z\\) is illegal. - Only a variable can be substituted e.g. \\(m/n\\) is illegal because \\(m\\) is an atom. The meaning of applying a term to a substitution is that every occurrence of \\(X_i\\) in the compound term is replaced with the corresponding substituent in \\(σ\\) simultaneously. This application is known as **instantiation**, and that new compound term \\(Eσ\\) is an **instance**. Bringing back the earlier rules about unification, you can see how unification is derived from substitution. Variables can unify with anything, so they are the term that is substituted in Prolog. There is an exception to this, which is known as the **occurs check**. We can use a substitution \\(σ\\) as a **unifier** for two terms if the application of \\(σ\\) to those terms makes them *syntactically equal*. This is distinct from *semantic* equality. When we talk about unifiers, we are only talking about the actual letters, not the meaning of the term. For \\(S = f(X,Y)\\) and \\(T=f(g(Z),Z)\\), if we have a \\(σ = \\{X/g(Z),Y/Z\\}\\) then \\(S\\) and \\(T\\) will be unified. Since the unification only cares about syntax, more than one unifier may exist for two terms. We could have just as well substituted the other way and it would still be a unifier. ## Unifier Computation ``` unify(X, Y, θ) = X = Xθ Y = Yθ case X is a variable that does not occur in Y:     return (θ{X/Y} ∪ {X/Y}) // this replaces X with Y in the unifier, and then adds a new substitution just in case X will appear later case Y is a variable that does not occur in X:     return (θ{Y/X} ∪ {Y/X}) // this replaces Y with X in the unifier. this is a rare case when the variables substitute each other case X and Y are identical constants     return Θ   case X and Y are compound terms like f(X1, ..., Xn) and f(Y1, ..., Yn)     return (fold_left (fun Θ (X,Y) -> unify(X, Y, Θ)) θ [(X1, Y1), ..., (Xn, Yn)]     // this applies the unification process to the compound terms     // the function is just folding left with θ as acc and the set of compound terms as the list ``` This is the basic pseudocode for the unify function. It's recursive, and side-effect free. However, it's not complete for writing an interpreter. For that, we also need backtracking. ## Backtracking The resolvent will maintain a list to satisfy our query. When we try to resolve a compound term, any goals within will be queued onto the resolvent to be resolved. This is *non-deterministic*, there is no order that will necessarily be followed other than that sub-goals within a single term will be inorder. ## Missionaries and Cannibals Let's imagine a problem where we have three missionaries and three cannibals that need to cross a river in one boat. If the cannibals outnumber the missionaries, they will get eaten. How can we get every person across the river? The concept of a *safe state* will be important here. A state is safe when no missionaries are eaten. By labelling certain states as unsafe, we can cut those paths out of our search algorithm. We should also define *transitions* between states, where a predicate moves a state from A to B. Our states are essentially a pair representing position. ```pro start(3-3-0-0-l). finish(0-0-3-3-_). ``` The elements in order are missionaries on the original side, cannibals on the original side, missionaries on the target side, cannibals on the other side, and the location of the boat. The safety of the state is based on whether there are more cannibals than missionaries. ```pro safe(0-_-M2-C2-_) :- M2 >= C2. safe(M1-C1-0-_-_) :- M1 >= C1. safe(M1-C1-M2-C2-_) :- M1 >= C1, M2 >= C2. ``` Every state here represents a point at which there are at least as many missionaries as cannibals on each side of the river. The change in state can be caused by a `carry/2` predicate which details how many missionaries and cannibals are being moved, respectively. In our case, the boat can only move up to 2 people so we need to represent every possible case of that. ```pro carry(2, 0). carry(1, 1). carry(0, 2). carry(1, 0). carry(0, 1). ``` However, we need a way to represent moving on both sides of the river, so we will have two transitions. ```prolog step(M1-C1-M2-C2-l, M3-C3-M4-C4-r) :- carry(X, Y), M1 >= X, M3 is M1 - X, M4 is M2 + X, C1 >= Y, C3 is C1 - Y, C4 is C2 + Y. step(M1-C1-M2-C2-r, M3-C3-M4-C4-l) :- carry(X, Y), M2 >= X, M4 is M2 - X, M3 is M1 + X, C2 >= Y, C4 is C2 - Y, C3 is C1 + Y. ``` This `step` predicate checks every valid `carry`, then will execute that transition from the current state. ```pro travel(A, A, _, []). travel(A, C, Visited, [B | Steps]) :- ``` ## Special Relativity - URL: https://sharifhsn.dev/blog/special-relativity/ - Structured data: https://sharifhsn.dev/api/posts/special-relativity.json - Description: Albert Einstein revolutionized the world of physics when he discovered the special theory of relativity. It is only possible due to his contributions that valuable technologies lik… - Date: 2022-04-05 - Exact published timestamp: 2022-04-05 - Topics: Physics - Categories: Physics - Source: Archive - Source URL: None Albert Einstein revolutionized the world of physics when he discovered the special theory of relativity. It is only possible due to his contributions that valuable technologies like GPS function. But what is it exactly? ## Dilation and Contraction The theory behind special relativity is two-fold: 1. All inertial frames of reference of equivalent. 2. Observers must measure the same value for the speed of light in a vacuum. This means that if even an observer has a huge difference in relative velocity, the speed of light is the same. This causes knock-on effects on the experience of time, and **time dilation**: $$ Δt = \frac{Δt_0}{\sqrt{1-\frac{v^2}{c^2}}} $$ The proper time \\(Δt_0\\) is the time interval with respect to an observer at rest. The dilated time is the time when the observer is at velocity \\(v\\). In the same way that time can be dilated, length can also be contracted: $$ L = L_0\sqrt{1-\frac{v^2}{c^2}} $$ The relationship here is inverse! Time can only be *dilated*, and length can only be *contracted*. If the velocity was the speed of light (impossible) then the time interval would be infinite, and the length of an object would be 0. ## Relativistic Addition Two objects traveling at different relativistic velocities further complicate these calculations. If we want to calculate the proper velocity \\(u\\) of an object from its relativistic velocity \\(u'\\) relative to an object at proper velocity \\(v\\) or vice versa, then we can use these equations: $$ u = \frac{u' + v}{1 + \frac{vu'}{c^2}} $$ $$ u' = \frac{u - v}{1 - \frac{vu}{c^2}} $$ As you can see, these are very similar equations, but their differences are important. ## Mass and Energy It's the moment you've been waiting for, the famous equation! $$ E = \frac{m_0c^2}{\sqrt{1 - \frac{v^2}{c^2}}} $$ Oh, it looks a little bit different. That's because just like length and time are relativistic, mass is as well. In order to calculate total energy from rest mass, we must apply the same expression to it. There's another way to calculate \\(E\\): $$ E = KE + E_0 $$ where \\(E_0\\) comes from the classic \\(E = mc^2\\) equation. We can also find the kinetic energy by rearranging these equations as $$ KE = (m - m_0)c^2 $$ from finding the difference between relativistic mass and proper mass. One more important attribute affected by relativity is momentum. As it depends on mass, it looks similar to the equation for \\(E\\): $$ p = \frac{m_0v}{\sqrt{1-\frac{v^2}{c^2}}} $$ ## Waves - URL: https://sharifhsn.dev/blog/waves/ - Structured data: https://sharifhsn.dev/api/posts/waves.json - Description: Light is curious because it is both a particle and a wave, which makes it behave strangely in certain contexts. When projected through a slit, some strange fringes will appear, whi… - Date: 2022-03-29 - Exact published timestamp: 2022-03-29 - Topics: Physics - Categories: Physics - Source: Archive - Source URL: None Light is curious because it is both a particle and a wave, which makes it behave strangely in certain contexts. When projected through a slit, some strange fringes will appear, which is the subject of the famous double slit experiment/ ## Interference When waves combine together, it is called **interference**. This interference can be either *constructive* or *destructive*. When light is projected through a slit, depending on the angle from the slit to the wall behind it, the light waves will either constructively interfere or destructively interfere, creating bright and dark fringes, respectively. The *position* of the bright fringes is given by the following equation: $$ d\sin{θ} = mλ ⇒ m = 0, 1, 2, 3... $$ where \\(d\\) is the distance *between the slits*, \\(θ\\) is the angle of interference from the central max onward, and \\(m\\) is the order of the interference fringe. For bright fringes, the order of the interference fringes begins at 0 for the central bright fringe, then increments onward for each surrounding fringe. Note that there is no functional difference between a fringe on either side of the central fringe as long as they are the same order away from it. Dark fringes have a similar equation: $$ d\sin{θ} = (m + \frac{1}{2})λ $$ The ½ is important here because the location alternates with the bright fringes. \\(m\\) is also different because it starts at the first dark fringe surrounding the bright fringe. You can think like this: ... **3** *2* **2** *1* **1** *0* **0** *0* **1** *1* **2** *2* **3** ... where the italicized numbers are dark fringes and the bold are bright fringes. ## Thin Film Interference has a curious corollary with refraction if done through a thin film. Light that is transmitted between two media is partially reflected and partially passes through if passing from a medium of lower to higher index of refraction e.g. air to water. It is fully reflected if from higher to lower. The partially reflected light (but *not* the fully reflected light) undergoes a phase shift by a half wavelength. Let's say we want to find the distance that the light travels in the medium, the "height" of the medium layer \\(t\\). If the light underwent a phase shift, the equation for interference is actually flipped: $$ 2t = (m + \frac{1}{2})λ_{film} $$ for constructive interference, and vice versa for destructive interference. However, if there is no net phase shift, then use the normal equations. Note that \\(λ_{film}\\) here is actually $$ λ_{film} = \frac{λ_{vacuum}}{n_{film}} $$ # ## Diffraction When a wave passes through a tiny slit, it spreads out. The slit has to be really small, as in on the order of the wavelength of the light, in order for this to happen, with more bending with a smaller slit. Diffraction has interference as well, so the equation for the position of dark fringes is similar, although different: $$ W\sin{θ} = mλ ⇒ m = 1,2,3... $$ where \\(W\\) is the width of the slit. **Importantly, \\(θ\\) is different here!!!** It is the angle between the center of the slit and the center of the dark fringe in question. \\(m\\) is also different in that it *cannot* be 0! The first dark fringe in diffraction has order 1, because at order 0 there is constructive interference. Diffraction grating has different properties than single slit diffraction, however, and is more similar to typical interference, following the constructive interference equation for bright fringes. Be careful about \\(d\\) in this context! It refers to the distance between each slit, which you might have to derive yourself from other values. ## The Link Layer - URL: https://sharifhsn.dev/blog/link-layer/ - Structured data: https://sharifhsn.dev/api/posts/link-layer.json - Description: Through all the layers of the internet, the most essential is the link layer: converting bits to and from signals, detecting errors, flow control, and addressing. - Date: 2022-03-28 - Exact published timestamp: 2022-03-28 - Topics: Internet Technology - Categories: Internet Technology - Source: Archive - Source URL: None Through all the layers of the internet, the most essential is the **link layer**: converting bits to and from signals, detecting errors, flow control, and addressing. ## Encoding We have spoken mostly through the guise of bits per second being sent across the internet and have left it at that. But what does it mean to send a bit over air waves? This isn't like a traditional electrical circuit or transistor in your computer; the internet uses signals. The nuances of digital signal processing are beyond these notes, so we will simplify here. We can imagine signals like a pseudo-transistor where low signals are 0 and high signals are 1. However, this causes two issues. One is that time synchronization is impossible for waves that travel at the speed of light, and the other is that because signals are not emitting all the time, we need a way to differentiate between signals and simple noise. The time synchronization is mandated by a clock standard such as the Manchester encoding, which is slower than the speed of light but allows each link to know exactly when a signal is coming. ## Framing A **frame** is a group of bits in sequence. Frames are useful for manipulating data at the link layer, but they can escalate small bit errors to big problems. If a frame is too big, then it will error out too often, so we will often have smaller frames. We also need a way to delineate the beginning and ending of frames. There are two ways, character stuffing and bit stuffing. With character stuffing, a special meta character will be used to delineate frames. `^` is used for the beginning of frame (BOF) and `$` is used for EOF. If you need a `$` in your text, then simply add an escape character, like another `$` right beforehand. With bit stuffing, a unique bit sequence `0x7E` or `0b01111110` will delineate frames. If that sequence appears in the data, it will insert a 0 within like `0b011111010`. ## Error Control Inevitably, bits will be corrupted on a physical link. This is a reality of circuit engineering. We have discussed several ways to handle errors on the application and protocol layer; what about on the link layer? We can either request retransmission as with TCP, or we can make corrections to the errors automatically, which may or may not be possible. Parity bits can ensure a fixed sequence so that the extra bit can protect against errors. For example, with even parity, there will always be an even number of 1s, so if a bit flip changes that, you know that there was an error. But multiple bit errors can escape this detection. Another method is when the sender will send a checksum computed from the message divided into chunks along with the message so that the receiver can use the checksum, just like with other layers. Polynomial codes utilize the high degree of uncertainty when computing `mod` of a polynomial. The remainder must be added to the message and it can be checked by the receiver. This is a more complex form of error detection mathematically. ## Addressing We have heretofore defined a host as an IP address. But how do we identifies hosts at the link layer? For links we use a **MAC address**, which is a unique 48 bit address to a device which will communicate through the internet. MAC addresses are far less sophisticated than IP addresses. They cannot be grouped or categorized like with NAT. The MAC address of the destination is included in the frame that is sent. They are also unique to types of links, so a device will have different MAC addresses for Wi-Fi, Bluetooth, etc. Hosts can then interact in multiple ways. One way is *access point*. Hosts must communicate with an access point which acts as an intermediary between hosts. Another is *ad hoc*. When you set up your Amazon Echo for the first time, you connect with it ad hoc via Bluetooth on your app and set it up from there to program it. ## ARP When connecting via LAN, hosts are part of the same subnet. In order to translate the IP address to a MAC address, the hosts use a protocol called **ARP (Address Resolution Protocol**. This is a simple protocol which defines its own identification and a source/destination address, as well as a broadcast `FF` packet. We know the sender IP, MAC, and the destination IP, but we don't know the destination MAC address, so that field is blank. The ARP packet will traverse the entire link, and if a host recognizes its own IP address, then it will send an ARP reply back to the source with its MAC address. ## LAN Extension LANs seem pretty nice. Why don't we just use LANs everywhere? Bandwidth is limited, so you need to limit the amount of hosts that can connect on one network. A *learning bridge* is a type of access point that connects LAN segments. It maintains a table of hosts that are connected to the network. It intermediates between the hosts, receiving and sending packets as necessary if packets are sent between LANs. For each packet, the bridge stores the MAC and port of the packet and then floods all the LANs if it can't find a match. ## Multiple Access If two packets are sent at the same time, that is called *collision* or *interference*. We can only have one pair communicating at a time. We must have a multiple access protocol to make sure that only one access method is used at a time. You can think of being in a class; if everyone spoke at once, it would be mayhem! There must be an extra overhead of each person raising their hand in order to ensure correct communication takes place. ## Wireless When a host sends a wireless signal that bounces off of a satellite to a receiver, how do we know that interference has occurred? We could use a TCP packet timeout, but that takes a while. A neat trick that you can do is leverage the fact that satellites reflect all signals to everyone, including your own. A host can receive its own signal and check if it was interfered with. Originally, **pure ALOHA** was used, based on the needs of a university in Hawaii. This uses the broadcast method. If it receives garbage back, it will sleep for a random amount of time and retransmit. The reason it is random is because if both the hosts that sent garbage data send the packets again at the same time, it will cause *another* collision! For transmission time \\(t\\), the vulnerable period where packets will be destroyed is \\(2t\\). **Slotted ALOHA** has better checks to stop collisions from occurring in the first place by setting slots where hosts will send packets, which means that they wouldn't collide in different slots. But the best throughput is done if we check the actual channel before transmitting the packet, so we don't have to guess or check after the fact if a collision happened. For this we need **CSMA (Carrier Sense Multiple Access)**. If the channel is sensed and there are packets travelling, the host will wait to send the packet then immediately send a packet when free. This is a persistent system. The problem with this is that this will cause a collision when multiple hosts are waiting for a channel. We can solve this the way we solved ALOHA. If each host sleeps a random amount of time, even a small amount of time, it will give the hosts time to sense the channel. ## Contention Access Methods In order to sense the channel, there are many methods. **1-Persistent CSMA** will constantly bug the channel to check if the channel is idle. If it's busy, transmit and check again. This only really works when there is one host trying to connect. **Non-Persistent CSMA** will transmit, then sleep for a random amount of time before trying again if the channel is busy. This means that that collisions won't happen. This is better than the previous method, but can be slow. **\\(P\\)-Persistent CSMA** has the same idea, but instead of being random it will transmit again with probability \\(P\\). This fine-tunes the previous method for our specific needs and improves utilization. Ethernet uses the **Ethernet Backoff Algorithm**, which utilizes binary exponential backoff. If a collision happens, then it will pick a time slot to transmit again out of \\(2^k\\) slots, where \\(k\\) is the number of collisions that have occurred. The length of each time slot is equivalent to. For example, let's say two hosts collide on an Ethernet connection. Both hosts will both, on their own, divide the next few seconds into two slots and randomly pick between them. If they each end up in different slots, they'll transmit fine. If they end up in the same slot again, repeat the process but with four slots. As you can see, the chance of collision will dramatically decrease with more slots. After a maximum of sixteen slots, if there are still collisions, then your host will give up and say that the link is down. ## Ethernet Ethernet is a wired multiple access protocol defined by IEEE 802.3. As previously stated, it uses 1-persistent CSMA with the backoff algorithm. Ethernet frames have a preamble full of special information, as well as the source/destination MAC addresses already mentioned. The header will also say the type of protocol being used, whether it is IPv4, IPv6, ARP, RARP, etc. Finally, at the end, a checksum will be stored so that we can check if corruption has occurred. | Preamble | Source MAC | Dest. MAC | Type | Data | Checksum | | -------- | ---------- | --------- | ---- | ---- | -------- | ## Multiple Access Channel Partitioning We have previously assumed that all hosts are fighting for the use of one channel and will have to retry connections if collisions occur. One way to solve this is to have predetermined allocation of channel access. This way, there is no chance of collision because each user will wait its turn on its own. However, this can introduce wastage if users are assigned to a channel partition but end up being idle. The division is governed by different techniques. **TDMA (Time Division Multiple Access)** will divide the spectrum across time. Like how OS schedulers will give processes exclusive access to the CPU for a certain amount of time, so too will users be granted exclusive access to the channel for a certain amount of time. **FDMA (Frequency Division Multiple Access)** will divide the spectrum across frequencies. The most common example of this is for radio. Radio channels are exclusive to certain frequencies. **CDMA (Code Division Multiple Access)** will send signals in a coded format. Imagine a large group of people that are speaking all at the same time. If everyone is speaking in English, the message will quickly get garbled. However, if every conversation has its own language, then the message will be clear. Even if I hear a conversation next to me in Chinese, I will tune it out because I don't know Chinese and be able to speak and hear in English. There must be some amount of power control so that one person does not speak too loud, but otherwise the communication will work. ## LAN Computers can be connected through a sort of wireless LAN. If connecting a computer that does not have a screen or an otherwise easily accessible user interface, it is often convenient to connect with it using ad-hoc mode to a device with a screen so it can be configured, like Amazon Echo. Each local set of computers that can connect in these ways is known as a **BSS (Basic Service Set)**. This is known commonly as a **LAN**. The sender will wait for the sense channel to be idle than transmit the entire frame, else do a random timeout multiplying by 2 as with Ethernet, and the access point will return an ACK. Access control in this sense is mediated by the time waited to check if the channel is busy or not. However, this can still cause collisions if a router does not hear from another, which is the **hidden terminal problem**. An **exposed terminal problem** may occur when a terminal is not properly signaled to not send because it is busy. In order to mediate these issues, senders send a small packet called **RTS (request to send)** and a receiver will send a small packet called **CTS (clear to send)**. If senders hear an RTS from somewhere else, they'll wait until the receiver is clear to send to. If there isn't a CTS, they can transmit because the receiver is open in that time. In this context, the HTP and ETP are caused when neither RTS nor CTS are heard, or when RTS is heard but not CTS. To avoid these issues, we can reserve channels in the same way that channels are reserved in CSMA. The small reservation packets are unlikely to collide compared to the large frame packets. ## Bluetooth **Bluetooth** is a short-range kind of Wi-Fi technology that operates in the ISM band of 2.4GHz to 2.8 GHz. It was initially created for use by wireless mice and keyboards. The data rate goes up to 721 Kbps, which is fine because it's not intended for data transfer, just communicating with nearby wireless devices. In order to code this message, **frequency hopping will be used**. A frequency sequence will be sent via CDMA about which frequencies will be used to transmit data in sequence. Then, the sender will send the sequence in parts at each frequency in order. The receiver which knows the frequency sequence will be able to tune into the right frequencies at the right time and receive the correct bitstream. Anyone else listening to the bitstream that doesn't know the frequency sequence won't understand it. Personal area networks under Bluetooth are either **Piconet** or **Scatternet**. Piconet has master/slave nodes, where the master is the one that allocates the Bluetooth channels among the slaves. Scatternet allows devices to be either master or slave in different networks so that a device can act differently in Bluetooth in different settings. Connection is established by beginning with an inquiry broadcast by the Bluetooth device. A potential slave will respond and the master will send back a page for the hopping sequence. The other device will establish itself as a slave through a slave response and the master device will do likewise, at which time the connection is established through ACK-DAC. ## Prolog - URL: https://sharifhsn.dev/blog/prolog/ - Structured data: https://sharifhsn.dev/api/posts/prolog.json - Description: Prolog is a special programming language that is not like many others. It exemplifies the paradigm of logic programming. - Date: 2022-03-24 - Exact published timestamp: 2022-03-24 - Topics: Principles of Programming Languages - Categories: Principles of Programming Languages - Source: Archive - Source URL: None **Prolog** is a special programming language that is not like many others. It exemplifies the paradigm of **logic programming**. ## Terms Functional programming is based on the idea that all operations exist using functions that give the same output for a given input. Imperative programming is based on the idea that we can manipulate variables and order them around like soldiers. Logic programming takes propositional logic from discrete structures and turns it into a language. Every program is made up of **terms**, which are **constants**, **variables**, and **compound terms**. A constant is any word that is lowercase, whether it be a number, a string, or just a word. A variable is a word that begins with a capital letter. Compound terms are relations. Relations in the logical sense are a subset of the Cartesian product of two sets, where the result set consists of elements that are contained in both. The top level name is the **functor**, and the number of arguments is the **arity**. For example, a compound term might be `male(robb)` where `male` is the functor with an arity of 1. **The order of arguments within the compound term matters!** **Facts** are a knowledge base for a particular program. Each fact is a compound term showing a relation, followed by a period. **Rules** are a generality on facts where they are facts *given* assumptions. A fact is just a rule that is true without assumptions. **The order of facts/rules is very important to the execution of code!** Prolog reads from top to bottom, so the most important facts/rules should be placed near the beginning. ```pro father(rickard, ned). father(rickard, brandon). father(rickard, lyanna). father(ned, robb). father(ned, sansa). father(ned, arya). ``` ## Queries Once we have our knowledge base, we can **query** for information. ```pro ?- father(ned, sansa). ``` This query will return true because it is in our list of facts. However, what happens if we query something it doesn't know? ```pro ?- father(ned, bran). ``` Prolog operates on the **closed world assumption**, which means that it only knows what it's been told. If it doesn't know something, it will assume falsity. So Prolog will say the Ned is *not* Bran's father, despite it having no idea whether that is the case. There are also **existential queries**. Instead of asking a boolean question, it asks if a fact exists. ```pro ?- father(ned, X). ``` This query asks for all the relations `father` where `ned` is the first argument. It will return ```pro X = robb ; X = sansa ; X = arya . ``` in the order that the facts are laid out. The semicolon is the logical OR in Prolog, so it's saying that either one of the results is true. It would return false if there are no facts that satisfies the query. ## Rules Rules in general can be any logical statements from facts that are inferred inductively. The general form is ```pro H :- B1, B2, B3, ..., BN ``` means that the head `H` is true if `B1` ∧ `B2` ∧ `B3` ... etc. ```pro parent(X, Y) :- father(X, Y). ``` means that the relation `parent` exists between `X` and `Y` as long as the relation `father` exists between `X` and `Y`. We can construct rules recursively as well: ```pro ancestor(X, Y) :- parent(X, Y) ancestor(X, Y) :- parent(X, Z), ancestor(Z, Y) ``` This defines ancestor in two ways. The first rule says that `X` is an ancestor of `Y` as long as `X` is a parent of `Y`. The second rule introduces an additional definition where `X` is an ancestor of `Y` as long as `X` is a parent of some `Z` which is an ancestor of `Y`. We can almost think of this like a list, where parent means adjacent and ancestor means inorder. Of course two elements are inorder if they are adjacent, that's the first rule. However, if we can imagine a long list, two elements are inorder if the adjacent element to `X` is inorder previous to `Y`! The logic applies in any general situation; this is the advantage of the logical paradigm. ## Unification The core of how Prolog will compute answers is through **unification**, which is similar to pattern matching. It has three unification conditions: - identical constants will be unified - variables will always unify - compound terms will unify if functor and arity match as well as their arguments, recursively (using the top two rules) ## Lists Data structures seem pretty incongruent with our idea of logical programming so far. Lists are defined similarly to OCaml as recursive data structures. ```pro [1, 2, 3, 4] % list [] % nil [H | T] % head cons tail ``` To find the last element of a list, we can use this rule ```pro last([H], H). last([_ | T], V) :- last(T, V). ``` The first rule is the base case, where if there is only one element in a list, that element is the "last element". Then, the second rule is recursively defined. If the first argument ## Arithmetic The arithmetic operators are built into Prolog, and are built as relations between terms. Arithmetic equality is *not* the same thing as unification. If you want to check for arithmetic equality, you must use `is`. The `is` operator evaluates the right hand expression and unifies the expression with the left. There can be variables present, but the variables need to be defined as some ground value. You can't create an algebraic equation using `is`. Basic arithmetic operators are also built in, but only operate on numerals. If you try to apply, for instance, `+`, to lowercase word constants, there will be an error. ## Backtracking Prolog computes the length of a list very similar to OCaml. An example OCaml length function might be: ```ocaml let rec len l = match l with | [] -> 0 | h :: t -> len t + 1 ``` The Prolog rules are similar: ```pro len([], 0). len([H | T], N) :- len(T, M), N is M + 1. ``` The relation `len` with arity 2 defined on the empty list maps to 0, obviously. The relation `len` with arity 2 and the first argument being a list with head and tail is the actual recursive part we are interested in. You might raise an eyebrow at the use of `is` in the second rule; after all, didn't I just say that the right hand side must have ground terms? We don't know what `M` is! Actually, we do know. You can think of the order of clauses as the imperative definition of variables. The variable `M` is declared and initialized with the first `len` clause, and it can then be used in the second clause. If we put the second clause first, Prolog would raise an error. You might notice a pattern here. The function is almost identical to the OCaml function, except for one difference. In the functional paradigm, every function evaluates to an expression which is returned within the function, and those are the values that are used for every operation. In the logical paradigm, there is no implicit return. `len(T, M)` is the same thing as `M = len(T)`! The length is computed using **backtracking**. It will try every possible rule for `len` and then evaluate every single one. You can think of a depth first search of the tree of rules. Every point at which there are multiple branches is called a **choice point**. In the `len` example, it can choose to either check the fact for `[]` or the recursive rule. Backtracking can cause unexpected results. Say we try to query ```pro ?- len(A, 2) ``` We're querying for a list of size 2 which *doesn't exist*. What would happen? Well, Prolog will try to find a matching branch to an existing rule. It will fail the first rule most of the time, so it will match the second rule. Because `M = N - 1`, when trying to find `M` it will subtract from the second argument and backtrack. It will eventually match to return the result `A = [h1, h2];` when it matches 0 as the length and backtracks to adding arbitrary variables to the list of T. But then something unexpected happens. After printing that last result, Prolog would spin on and on, never returning again. That's because although we have matched the first rule and returned a list, Prolog will try every rule to see if it matches. And in fact, the second rule does match `[T2, 0]` because a variable will match any value. Prolog will endlessly backtrack because a choice will always be available that can match the new recursion. This is much different from OCaml which will only execute the first match arm that matches. You need to be careful in Prolog in order to prevent this kind of result. ## Tail Recursion Our `len` function is not tail recursive, since it adds on to the recursive call. In order to optimize our code, including an accumulator will aid because of tail recursion which we have discussed before. ```pro len2([], Acc, Acc). len2([H | T], Acc, N) :- M is Acc + 1, len2(T, M, N). ``` What we've done here is shift the `is` clause before the recursive call, which means we don't need to hold on to that choice when recursively calling `len2`, since it's guaranteed to already be checked. You might ask, why do we even need to have that `is` clause? We might as well directly call `len2(T, Acc + 1, N)`. However, there's an issue here. The `len2` relation only *unifies* terms, which is a powerful operation, but it *does not evaluate them*. For a list `[1, 5, 7, 3]` it would show `N = 0 + 1 + 1 + 1 + 1`, not `N = 4` which would be correct. In order to make this correct, we need to evaluate the addition before making the recursive call into the second clause because that clause does not perform any evaluation. This means that when we call `len2` we need to do `len2([], 0, X)`. This is a little annoying; we always know that our accumulator will begin at 0, why do we need to include it? There is a way to elide this definition: ```pro len2([) ``` ## Append Another instructive example is the simple list append. ```ocaml let rec append p q: match p with | [] -> q | [h :: t] -> h :: (append t q) ``` ```pro append([], Q, Q). append([H | P], Q, [H | R]) :- append(P, Q, R). ``` ## Prefix/Suffix For a given list, we can place a separator in the list. Every element prior to the separator is the **prefix** and every element after is the **suffix**. The separator can be anywhere, including before or after every element. We can get all possible prefixes/suffixes for a list for all separators: ```pro ?- prefix(X, [1, 2, 3]). ``` ```bash X = []; X = [1]; X = [1, 2]; X = [1, 2, 3]. ``` Let's think about how to build this out logically. We need to find all lists which are possible prefixes of this list. That means we don't want to capture, for example, `[2, 3]`, because that is not a possible prefix. We have to start from the left and include from there. ```pro prefix([], []). ``` The empty list should map onto the empty list, naturally ## Generate and Test We need a general approach to solve problems from a Prolog standpoint. The paradigm is very different from imperative or functional, so we can't just look at it the way that that OCaml or Java does. Let's make `take`, which will remove exactly one element `x` from a list. ```pro take([H | T], H, T). take([H | T], R, [H | S]) :- take(T, R, S). ``` The first rule is the most obvious. If the element is at the head of the list, we can just return the tail of the list. This is our base case. The second rule maps onto the case in the middle of the list. To imagine this, let's use an example ```pro ?- take([1, 2, 3], 2, T). ``` We try to unify this to the first rule, but it fails because 2 is not at the head of the first list. The second rule will unify because the first element is a list with multiple values, the second is a variable that will unify anything, and the third element is a variable that will unify anything. Upon match, it will query `take([2, 3], 2, S)`. This query will unify with the first rule, which will replace `S` with the value `[ 3 ]`. Once that `S` is unified, the rule will backtrack to the initial query which asked for `[H | S]`. Well, `H` in that initial case was `1` so our final return `T` will be `[1, 3]`. By extending `take`, we can implement other functions like permutations. ```pro perm([], []). perm(L, [H | T]) :- take(L, H, R), perm(R, T). ``` Let's permute `[1, 2, 3]` once again. ```ocaml ?- perm([1, 2, 3], X). ``` It will not unify with the first rule, but it will with the second, obviously. The second clause will permute the list with the head taken away. Now, you might notice that `H` is not an explicit value here as when we used it earlier. Prolog seeing this will try to unify it with *every* value in the list, so `take` will be queried three times, with each value in the list being used in `H`. The resulting `R` from each of those queries is permuted again. The `H` used in `take` will be the head that is cons the backtracked form. `[1, 2, 3]` will become `take([1, 2, 3], H, R)`. Prolog will run through every choice, starting with `take([ 1,[1, 2, 3] R)` which will unify `R` with `[2, 3]`. Now, the `perm` clause will be `perm([2, 3], T)`. Let's assume that this gives all valid permutations of `[2, 3]` i.e. `[2, 3]` and `[3, 2]`. Well, the resulting match on the original `X` would be `H | [2, 3]` and `H | [3, 2]` or for this particular unification, `[1, 2, 3]` and `[1, 3, 2]`. This conclusively shows the inductive step, and the base case is trivially the empty list. As you might be able to tell from these examples, there is a general strategy to solving problems in Prolog. 1. **Generate a solution.** 2. **Test if it is valid.** 3. **If not valid, backtrack and try another solution.** ## Quicksort One way to sort a list in Prolog is to use the generate and test method. You can generate all permutations of a list, and test whether each list is sorted. This is obviously incredibly inefficient with a Big O of \\(O(n!)\\). ```pro partition([], Y, [], []). partition([X|Xs], Y, {X|Ls], Rs) :- X =< Y, partition(Xs, Y, Ls, Rs). partition([X|Xs], Y, Ls, [X|Rs]) :- X > Y, partition(Xs, Y, Ls, Rs). quicksort([H | T], SL) :- partition(T, H, Ls, Rs), quicksort(Ls, SLs), quicksort (Rs, SRs), append(SLs, [H|SRs], SL). quicksort([], []). ``` This quicksort method is a very declarative manifestation of how to do quicksort in plain English. Partition the list along the pivot, quicksort the left and right, then combine them. ## 8-Queens One classic problem in math is to find all of the arrangements of queens on a chessboard where none of the queens threaten each other. We will find it for 8 queens on an 8×8 chessboard. Let's cut down on our sample space, because there are *a lot* of permutations. If a queen is next to another, it obviously threatens it. In fact, because queens can never share the same row or column, that massively cuts our sample space. We can describe our solution as an array of eight values from one through eight, where the index of the value is the row of the queen and the value itself being its column. This way of describing the sample space not only cuts our sample space but makes our computation simpler. However, we have not checked the diagonals. That is what we need Prolog for. ```pro checkBoard([H | T]) :- L is H - 1, R is H + 1, checkRow(T, L, R), checkBoard(T). checkRow([H | T], L, R) :- H =\= L, H =\= R, LN is L - 1, RN is R + 1, checkRow(T, LN, RN) checkBoard([]). checkRow([], _, _). ``` ## Threads - URL: https://sharifhsn.dev/blog/threads/ - Structured data: https://sharifhsn.dev/api/posts/threads.json - Description: Concurrency is a powerful tool to increase performance in applications. This is usually accomplished using threads. - Date: 2022-03-23 - Exact published timestamp: 2022-03-23 - Topics: Operating Systems Design - Categories: Operating Systems Design - Source: Archive - Source URL: None Concurrency is a powerful tool to increase performance in applications. This is usually accomplished using threads. ## Motivation Clock frequency in CPUs has massively increased over the past few decades, growing at an exponential rate from the tens of MHz to GHz. However, circa 2005, clock frequency began to stagnate, and for the past decade the clock frequency of consumer CPUs have not increased past around 4 GHz. Why is that? Clock frequency is typically increased by increasing the number of transistors in a single CPU. This can be done either by increasing the size of the CPU, or decreasing the size of the transistors. Because we don't want huge CPUs in our computers emitting tons of heat, we have opted to decrease the size of transistors. However, there is a limit to how small they can go. When a transistor starts getting smaller than 30 nm, power leakage will begin to affect the CPU, causing the same heat problems that a big CPU does. In order to keep up with computing trends, new CPUs have multiple CPU cores with the same transistor density. However, utilizing multiple CPU cores requires parallelism in the CPU. An ideal CPU might run one process on each core, with each process getting full control over the CPU. You can visualize this idea like a highway. Having multiple lanes allows cars to travel much faster than if there is only one lane. However, there can be issues when the lanes must combine, like in merging. This causes traffic like if there was only one lane, and can actually increase overhead as cars attempt to merge. Decreasing the amount of merging (or **synchronization**) is key to fast concurrency. ## Threads We have discussed processes before; **threads** are not much different. They contain information about an execution such as the register values, open file streams, etc., but are more lightweight than processes. There are multiple frameworks for threading that are useful for different workflows. The **producer/consumer** model is typically used for data analysis. Multiple threads create some data which is handled by a synchronizing consumer thread. The **pipeline** model works like a real pipeline, where multiple threads are consumed in steps and pipeline flows. This is common in GPU programming. The **background** model will work similarly to the interactive/batch model in MLFQ. Background threads do behind-the-scenes dirty work when the CPU is idle, and foreground threads will run the actual important parts of a process. Since threads exist within a single process, they share the same address space, including the stack and the heap. This can cause many, many bugs when it comes to accessing and modifying memory. We will discuss methods to solve these bugs later. Threads do not share the same instruction pointer, however. Each thread can execute from different parts of the program. All register values are unique to each thread, as they can execute on any CPU. Threads can technically share the same stack on the address space. However, this is pretty obviously a bad idea. Typically, threads will create their own stack so that each thread does not interfere with each other's local variables and execution, with their own stack pointer `%esp` to point to their stack. ## OS Support It's all well and good to have threads, but we need a way for the OS to know how we're switching between threads. There are multiple ways to do this. **Userspace threads** are a way to have concurrency regardless of what the OS thinks. A run-time will execute along with your program and swap between threads itself without the OS knowing. This is how `pthread` works, and also how Project 2 does. However, the power that concurrency grants for multicore CPUs is still dependent on the number of kernel threads generated, because those are the only ones that run on each CPU. **Kernel threads** can be associated one-to-one with user-level threads. This way, threads can take advantage of multiprocessing because the kernel will parallelize operations across CPUs. However, this incurs an overhead for the syscalls needed to parallelize threads, and there will never be enough kernel threads for the user threads. ## Scheduling Small operations like incrementing a variable can be difficult when we have threads accessing the same data. The simple instruction `balance++` where `balance` is stored at memory location `0x9cd4` would have this assembly output: ```nasm 0x195 mov 0x9cd4, %eax 0x19a add $0x1, %eax 0x19d    mov %eax, 0x9cd4 ``` If we have a context switch when `%eip` is at `0x19a`, the first thread will `add` but not store the new value back in memory. That means that the variable `balance` would be less than expected because the `add` by the first thread would not be stored into memory. This can lead to *non-determinism*, where the same input will cause different outputs, as well as *race conditions*, where the result of a program will depend on the CPU timing of different operations. We want these three instructions to be executed uninterrupted. This is known as **atomicity**. In this **critical section**, no process or thread can interrupt the currently executing thread. However, if we allow programs to do this, then they can take advantage of this property and essentially disable the timer for context switches by declaring the entirety of execution atomic. We can't have full blocking on interrupts for threads, so we only lock other threads from executing at the critical section. When the timer interrupt hits, the scheduler will still take control, but no other thread will be able to execute at the critical section. ## Synchronization Primitives There are many high-level **synchronization primitives** in the OS that exist to ensure the correct order of instructions. Each kind of lock is designed to solve a specific problem; no single lock can solve all of them. We want to have correctness in our concurrency. This means that we must have mutual exclusion where only one thread can access the critical section at a time. If multiple threads are waiting for something, they cannot be stuck forever, they must make progress. Similarly, threads cannot be forced to wait an unreasonably large amount of time. Concurrency must also be fair and not favor any thread over any thread. And obviously, we do not want to overly tax the CPU with the overhead of synchronization. We need underlying hardware atomic operations in order to implement synchronization. When these hardware instructions execute, *no instruction or interrupt* can execute. Example instructions are `Test&Set` and `Compare&Swap`. The reason we need this is because if we don't make these instructions atomic, we can encounter race conditions. Imagine an unset lock with a while loop and two threads. The first thread successfully enters the loop because the lock is unacquired. However, before it can acquire the lock, the thread switches and the second thread executes and acquires the lock. Now the first thread has not acquired the lock but it is still in the loop, so when it runs again it will acquire the lock even though it has already been acquired, allowing two threads in the same critical section! In order to combat this, testing and setting are the same atomic operation so there isn't a gap that can allow a switch in between those two operations. ## Spinlocks Making our implementation fair is not as easy. If we imagine the most basic spinlock which switches back and forth, a program that knows how long a context switch will take will acquire the lock right before every context switch so no other program will be obtain it. In order to introduce fairness, the ticket system for schedulers can be used here. Every thread is assigned a ticket for every turn it gets, with the first thread getting ticket 1, second ticket 2, etc. Now when a thread releases a lock, the ticket increments to the next thread. Even if it wants to maliciously acquire the lock before a context switch occurs, it can't because it has lost its turn. We need a special atomic function to make this work like `Test&Set` with basic mutex. We should be able to get a ticket number and increment it atomically. The function used for this is `Fetch&Add`. ```c int FetchAndAdd(int *ptr) { int old = *ptr; // these two lines are *ptr = old + 1; // executed atomically! return old; } ``` Spinlocks are fast when we have a short critical section because we're not context switching very often, so when we switch between locks constantly we want to avoid that overhead as much as possible as that dominates execution time. However, in a situation where there is only one CPU, letting the threads spin on a lock is extremely wasteful. The CPU scheduler has no idea that a thread is waiting for a lock so it will ignorantly run it even though all it does is spin. One alternative is to yield the thread instead of letting it spin, telling the CPU scheduler that the thread doesn't need to run right now. *In general*, a shorter critical section is okay to spin, but a longer critical section is not okay to spin. ## Condition Variables Mutex locks are just a way of restricting access to a critical section. It's synchronous, but there's no order to the way that the threads have to run. That is up to the caprices of the scheduler. However, let's say that ordering is significant to the execution of our threads. How can we ensure that threads are executed in a certain order? We need some functions and concepts to implement this. There are two system call functions `wait(cond_t, mutex_t)` and `signal(cond_t)` that are used with **condition variables** `cond_t`. Let's assume that a thread has acquired the lock, but it has not satisfied the condition to run yet because another thread must run first to preserve ordering. In that case, the condition variable will indicate that the thread cannot run and it will call `wait` to sleep the thread until the condition variable changes, releasing the lock so that the other thread can run. The caller of `signal` wants to wake up that sleeping thread, unless there are no sleeping threads, in which case it does nothing. The best time to use condition variables is in a producer/consumer system, like in a pipe. Producers write to a pipe, and consumers read from the pipe. The pipe buffer has a limited size, so the producer can only add data to it if the buffer is empty, and the consumer can only read from it if the buffer has contents. For simplicity, assume a buffer of single unit size for now. ```c void *producer(void *arg) { for (Int i = 0; i < loops; i++) { Mutex_lock(&m); // lock the mutex while (numfull == max) { // while the buffer is full, Cond_wait(&cond, &m;) // wait } do_fill(i); // put data in buffer Cond_signal(&cond); // signal conditional variable Mutex_unlock(&m); // unlock the mutex } } void *consumer(void *arg) { while(1) { // consume on and on forever Mutex_lock(&m); // lock the mutex while (numfull == 0) { // while buffer is empty, Cond_wait(&cond, &m); // wait } int tmp = do_get(); // get the data Cond_signal(&cond); // signal conditional variable Mutex_unlock(&m); // unlock the mutex printf("%d\n", tmp); // do something with data } } ``` Let's consider an example with a producer and two consumers, as well as a FIFO scheduler. The first consumer thread runs and waits immediately because the buffer is empty. The second consumer thread does the same. The producer thread runs through its entire loop, filling the buffer, then when it loops again it will wait (since our buffer is size 1). The first consumer thread will return to the while loop, and now that the buffer is no longer empty it will fetch the data, signal the thread, and finish. However, there is a problem here. The `Cond_signal` function signal is extremely generic and does not signal a specific thread. Instead, it will signal a random thread, which may be the producer or other consumer thread. But the whole point of using conditional variables is so that we have ordering! We need to have *multiple* conditional variables: ```c,hl_lines=5 void *producer(void *arg) { for (Int i = 0; i < loops; i++) { Mutex_lock(&m); // lock the mutex while (numfull == max) { // while the buffer is full, Cond_wait(&empty, &m;) // wait until buffer is empty } do_fill(i); // put data in buffer Cond_signal(&fill); // signal to show buffer is somewhat filled Mutex_unlock(&m); // unlock the mutex } } void *consumer(void *arg) { while(1) { // consume on and on forever Mutex_lock(&m); // lock the mutex while (numfull == 0) { // while buffer is empty, Cond_wait(&fill, &m); // wait until buffer is somewhat filled } int tmp = do_get(); // get the data Cond_signal(&empty); // signal to show that buffer is empty Mutex_unlock(&m); // unlock the mutex printf("%d\n", tmp); // do something with data } } ``` In this example, the consumer thread will only run when the producer has signaled that there is content in the buffer. ## Semaphore Conditional variables do not have any state. They only signal the condition to wake or sleep a thread. This can be useful if you have lots of threads that do the exact same thing, like in a producer-consumer model. However, if you want a more complex way to represent state, you need a **semaphore**. Semaphores are essentially conditional variables that contain an integer value which can be changed by threads. This allows threads to have multiple options based on a single variable shared across multiple threads just by changing the value in the semaphore. The semaphore is incremented every time a thread is woken, and decremented every time a thread sleeps. The atomic operations for semaphores are `Allocate&Initialize`, `Wait`, and `Post`. We must first create the semaphore by allocating/initializing it. The `Wait` operation decrements the semaphore if it is greater than 0, and `Post` will increment the semaphore and wake a sleeping thread. In this context, joining and exiting threads have no complication with mutexes; all you need to do is wait and post the semaphore, respectively. We can actually construct locks using semaphores. Acquiring and releasing locks work the same way as waiting and posting, so you can just use those same atomic operations instead of `Test&Set`. In reverse, you can build a semaphore using locks and conditional variables, though in practice that's a little more complicated. ## Routing Algorithms - URL: https://sharifhsn.dev/blog/routing-algorithms/ - Structured data: https://sharifhsn.dev/api/posts/routing-algorithms.json - Description: When a host needs to find the best destination/prefix to another host, what routing algorithm does it use? We can use either intra or inter domain. - Date: 2022-03-21 - Exact published timestamp: 2022-03-21 - Topics: Internet Technology - Categories: Internet Technology - Source: Archive - Source URL: None When a host needs to find the best destination/prefix to another host, what **routing algorithm** does it use? We can use either *intra* or *inter* domain. ## Routing and Forwarding As you may recall, the *routing table* is a large DRAM table containing all possible paths and the *forwarding table* is a small SRAM table containing only the best possible path. The router also caches previous best paths so that they don't need to be retrieved for the same host over and over from DRAM. You can visualize multiple routers in an undirected weighted graph. The weights matter so that a slow link will not be chosen even if the path is technically slower. By maintaining all paths, the router can pick a new best path if a link in the route crashes. But how does the router actually update the forwarding table? There are several different algorithms we can use. ## Link State **Link state** is an intra-domain method that works fine for 50-100 routers, say on a university campus. Let's use the graph analogy again. It is often useful to abstract concepts in networking to graphs when thinking about algorithms, since we can apply well-known concepts from mathematics and algorithm design to the problem. There are multiple factors that play into the weight of each link, not just link speed, but routers assign their own holistic score to links for simplicity. The obvious choice here is to use [Dijkstra's algorithm](https://en.wikipedia.org/wiki/Dijkstra%27s_algorithm) for shortest path, as it applies to undirected weighted graphs like the on we have just constructed. This algorithm generates the shortest path to every path up to and including the destination. It requires knowing all of the link speeds beforehand. This is accomplished through a *link state broadcast* i.e. pinging every router and checking the link speed of the ping. The algorithm is also iterative, which means the number of shortest destinations scales with iterations linearly. The time complexity of the algorithm is \\(Θ(|E| + |V| \log |V|)\\). This algorithm works best in a centralized graph like at a university. Instead of sending a link state broadcast every time a connection is made, each node in a graph will send its own **link state** to a centralized location, and that location will be checked for recomputation of shortest path. The centralized node is the one that does all the calculations for shortest path. This is known as a **distance vector**. Each router must calculate its own distance on top of the other links to the node that it must pass by. Think of a table that is passed from link to link which is updated according to Dijkstra's algorithm. The initial distance vector table for a router only contains the link speeds for its adjacent routers, and all others are marked as infinity. Degenerate longer paths are thrown out when tables "merge" at a router. If a link breaks, then the routers that have broken links get reset in the final table. When the broken router sends a new distance vector table, it updates the central table in the same process as the table generation. However, this doesn't work well when we have different domains in different autonomous systems controlled by different ISPs. We need some kind of inter-AS algorithm to manage this. However, as the intra-AS algorithms works well, we use *federation* in our algorithms. Inter-AS algorithms are used to route between ASes and intra-AS algorithms are then used within the AS. ## BGP There are some barriers to inter-AS routing. We might not know whether the host is reachable, so we need to ask nearby ASes if they know. The protocol that these ASes use is **BGP**, the **Border Gateway Protocol**. With BGP, ASes can - obtain subnet reachability information from neighboring ASes - propagate reachability information to all the routers within the AS - use this information to calculate the best route to another subnet BGP works through persistent sessions between BGP routers over semi-permanent TCP connections. The sessions can be through any router, there doesn't need to be a physical link. The two peers can communicate regardless of the physical router. This is the sequence of steps: 1. Establish TCP over port 179. 2. Exchange routing tables. At first, this is a lot of data, but after a connection is established only deltas (like Git) need to be exchanged. 3. Send four kinds of messages. - *Open* - establish session by exchanging AS numbers and the BGP identifier (arbitrary router IP address). There is a timer for how long to wait before just assuming the session is down i.e. in an outage. - *Notification* - report unusual conditions, usually an error. If there are header errors, or a timer has expired, etc. then the TCP session must be terminated. - *Update* - inform if there are new active or old inactive routers. The message includes the withdrawn routes, which are inactive routers, the length of that field as it can be variable, etc. - *Keepalive* - regular message that the session is alive. BGP needs this regular message otherwise the neighbor AS won't know whether you're alive or not. When we say that a router is advertising a prefix, there are some necessary assumptions. The routing information must be valid, and does not need to be refreshed until necessary. The path that an AS advertises to a node is the same path that AS uses to communicate with that node. The BGP protocol additionally has many attributes that indicate the characteristics of a prefix: 1. ORIGIN: advertises the origin of an announcement and prefix injection. This is manually configured by intra-routing protocols. 2. AS_PATH: a link of ASes which is the path of ASes that the prefix has traveled to. This is useful for detecting loops and selecting shorter routes of ASes. 3. NEXT_HOP: the hop field is when you cross the AS boundary. After you cross, the next hop value is the next AS you need to get to in order to get to the destination router. The reason this is important is because the IP packets within an AS can travel in any order irrespective of BGP, so the intra-routing algorithm can decide where to go based on this value. This becomes significant when a network experiences *transit traffic* and different networks are using that network as an intermediary. In order to get the packet to exit the network as quickly as possible, BGP can make the decision to send the packet through non-BGP routers if it's a faster way to make the packet leave the network. 4. MED: stands for Multi-Exit Discriminator. If ASes are connected via multiple links, the AS that receives a prefix uses this value to discriminate between exits, with a lower value being better. 5. LOCAL_PREF: indicates a local preference for a specific prefix, is a number that is more for more preference, by default 100. For example, if outbound traffic is preferred to be a specific exit point, then that point will have high local preference. It can be used to break a tie between routers. ## Swapping - URL: https://sharifhsn.dev/blog/swapping/ - Structured data: https://sharifhsn.dev/api/posts/swapping.json - Description: Sometimes we don't have a lot of memory to access. In those cases, we need a place to put our data. The natural solution is to put data in the disk, which has very high storage. Th… - Date: 2022-03-21 - Exact published timestamp: 2022-03-21 - Topics: Operating Systems Design - Categories: Operating Systems Design - Source: Archive - Source URL: None Sometimes we don't have a lot of memory to access. In those cases, we need a place to put our data. The natural solution is to put data in the disk, which has very high storage. This is known as **swapping**. ## Mechanism The actual process of copying memory to disks is not complicated, it is simply file I/O and is supported by hardware. We swap memory out to disk, and swap in data from disk to memory. The complicated part comes when we introduce paging to this idea. In order to accommodate this, we add a *present bit* to our page table entry. If the bit is 0, then the page has been swapped out to disk and the physical frame is no longer valid. When there is an attempted translation of a page table entry with an unset present bit, a trap will occur and the hardware will swap in memory from the disk, setting the present bit with a new translation for the new physical frame number. An important note is that a process will *never* directly access the disk! Instead, a trap will swap disk and memory and the process will access the memory as dictated by hardware and the OS. ## AMAT Swapping can be slow. We want to improve our **Average Memory Access Time (AMAT)**. A hit in this context is an access that goes directly to RAM, and a miss is a trap that will get the memory from disk into RAM. $$AMAT = (Hit \\% \cdot T_m) + (Miss \\% \cdot T_d)$$ ## Policies When we swap in and out of memory, we need a way to find the most frequently accessed pages to reduce AMAT. There are several policies to accomplish this. The optimal policy which has oracular knowledge of future page accesses can keep pages that will be accessed sooner and replace pages that will be accessed later. If we know that a page will be required soon, we're definitely not going to swap it out. Alternatively, if we know that it's going to be a while before the page is going to be accessed, we can swap it out, improving our AMAT. This is obviously infeasible when page access is random, which is almost all the time. The simplest possible policy is, as with schedulers, **FIFO**. However, this does not work well. A better, also simple policy is **LRU**, or **Least Recently Used**. Most caches of this kind use LRU because of how effective it is despite its simplicity. The idea can be expressed in one line: swap out the page which has been least recently accessed. By keeping the pages of physical memory in a queue when they are created, it's simple to dequeue and swap out the page for a new, freshly enqueued page. The counter is reset every time that a page is swapped out. LRU adds more memory usage in order to maintain the queue > You would think that having more pages reduces the number of page misses. However, due to **Belady's Anomaly**, this is not always the case for FIFO. LRU does not suffer from this phenomenon However, LRU does not take into consideration how many times a page has been accessed. A large I/O scan that is only used once might flush memory because it is recent without regard to how often it is used. When you stream a movie, you are sequentially accessing that data. This is a common access pattern in computer science. We can use **pure LFU**, or **Least Frequently Used** to combat this. However, this runs into the opposite problem: what if a page was very frequently accessed in the past? The policy will not forget that so it won't be evicted despite it being very old. A better approach is to combine both of the ideas of recency and frequency, through algorithms such as **LRU-K** or **2Q**. However, these algorithms can be expensive due to their increased complexity compared to these other simpler algorithms. ## Implementing LRU There are two approaches: software and hardware, as always. A software based LRU will be a sorted linked list. Accessing a linked list is slow, which means \\(O(n)\\) complexity for a memory reference, although replacing pages is very fast at \\(O(1)\\). ## Routers - URL: https://sharifhsn.dev/blog/routers/ - Structured data: https://sharifhsn.dev/api/posts/routers.json - Description: Everyone nowadays gets the Internet through the router in their home. But what is a router? - Date: 2022-03-10 - Exact published timestamp: 2022-03-10 - Topics: Internet Technology - Categories: Internet Technology - Source: Archive - Source URL: None Everyone nowadays gets the Internet through the **router** in their home. But what *is* a router? ## Speed It's likely that you have an internet speed in the tens of Mbps. This unit is **megabits per second**, or one million bits per second. Some routers in important locations have speeds in the Gbps, which is 1 *billion* bits per second. This is incredibly quick. How do routers accomplish this? ## Structure The router is composed of a **control plane** and a **data path**. The control plane contains all the necessary external tooling to determine the paths. It contains routing protocols that need to be implemented, as well as a routing table. This table is used to look up all the possible paths to the destination and pick the shortest one. This information is all stored on the data path, where each packet is actually processed. The routing table shortest path is put into the **forwarding table** which typically has a very small amount of SRAM in order to quickly mediate routing. You can think of it almost like a register that is close the process as opposed to the main memory where the calculations are made. > Remember this distinction! The *routing table* contains all possible routes from this router, kind of like Google Maps. It needs to be a big table, so it is usually stored in slow, inexpensive DRAM. The *forwarding table* only contains the best possible route, so it is a small table stored in fast, expensive SRAM. Routers must also have input and output buffers. The input buffer is needed because the router may take in packets faster than it can process them, so it needs a place to put them before it can figure them out. The output buffer is needed because of the opposite problem; if the router processes the packets faster than the output link can output them, the packets need a place to be. ## Forwarding Engine Let's examine the forwarding table in a little more detail. The forwarding table is a simple table that looks like this: | Dest-network | Port | | ------------- | ---- | | 65.0.0.0/8 | 3 | | 128.9.0.0/16 | 1 | | 149.12.0.0/19 | 7 | It is a simple path from networks, remembering which port to send packets. You might have noticed that our destination networks are in prefix form. The one on port 7 has the longest prefix because it is matching the first 24 bits. We must match the longest prefix form first when looking for the network. These prefixes are the most specific and therefore must be matched before the smaller, more general prefixes. Prefixes can overlap, so we must pick the more specific one. It is typically not explicitly linear, a hash table-like structure is used to make lookups faster. In fact, **hierarchical addressing** is important for allowing us to aggregate routes. Different organizations within the same ISP network will have some bits after the network bits devoted to differentiating between organizations. Let's imagine that Rutgers Newark and NJIT are on the same ISP network, but are different organizations, obviously. The first 19 bits of that prefix might be devoted to the ISP network, then the next 4 bits will be used for organizations within that network. ## Switching Fabrics Let's consider each port as form of input and output. We will need a way to switch between devices. There are multiple ways to implement **switching fabrics**. One way is through **memory**. This is the cheapest way of going about it. The I/O device will take the network and write it into memory. We can use DMA (Direct Memory Access) to do this more quickly. The header is processed, and the router checks the routing table for the best route. The packet must be copied from the port into memory, then processed, then loaded back into the output port. This is very slow. Most of the bottleneck is in the memory bus when copying memory. These routers are typically less than 10 Mbps. Another way is through **bus**. The packet is transferred between input and output through a I/O bus, with a memory buffer to hold the packet on each end. There's no CPU movement here, it's all I/O instructions. But we still need some kind of command for the destination port number. The CPU will periodically calculate and *cache* the forwarding table on the input device (also known as **line card**). If there is a cache miss, then I/O will ask the CPU to update the cache from main memory, which is slow but doesn't happen often. However, since there's only one bus, there is only one communication possible between one port and another at a time. This speed is in the high Mbps. The third way, which is the most high-end way, is the **crossbar**. It works like a bus, except that there exists a bus between every single port, which means that for \\(n\\) ports, we can have \\(\frac{n}{2}\\) simultanenous transfers. This is in the best-case scenario where each input goes to a different output. However, if two ports are trying to connect to the same port, then one of them will be bottlenecked because it's still one bus. However, this complicated switching fabric is extremely expensive, which scales with the number of ports. The price can be worth it for a company like Google; their speeds can be in the Tbps! ## Mirrors - URL: https://sharifhsn.dev/blog/mirrors/ - Structured data: https://sharifhsn.dev/api/posts/mirrors.json - Description: Mirrors can interact with light in interesting ways. In fact, reflection and refraction are probably most people's engagement with laws of physics surrounding light in their daily … - Date: 2022-03-08 - Exact published timestamp: 2022-03-08 - Topics: Physics - Categories: Physics - Source: Archive - Source URL: None Mirrors can interact with light in interesting ways. In fact, reflection and refraction are probably most people's engagement with laws of physics surrounding light in their daily life. ## Reflection On a plane mirror, like the kind that you're used to in daily life, *the angle of incidence equals the angle of reflection*. To clarify some terms: *incidence* refers to the ray that is going towards the mirror, and *reflection* refers to the ray that is coming out of the mirror. The *normal* is a perpendicular line, in this case to the plane of the mirror. Because of this characteristic of plane mirrors, the reflection of an object (or **image**) will be upright and the same size as the object. The image will be "inside" the mirror, which is called *virtual*. The distance "within" the mirror to the image is the same distance from the mirror to the object. This is useful for real life because we want to see most objects the same as they actually exist. However, in physics, sometimes we want to manipulate these variables, and for that we use **spherical mirrors**. A spherical mirror is a mirror that is curved such that it could be cut out of a sphere. It can be either concave or convex (front and back of a spoon). There are two points that are significant to a mirror, the focal point \\(F\\) and the center of curvature \\(C\\). If the mirror was part of a sphere, then the center of curvature would be the center of that sphere. The focal point is where parallel rays intersect after reflection. Both points are on the *principal axis* of the mirror. The focal point is halfway between \\(C\\) and the center of the mirror. In order to determine how the image of an object looks when reflected off of a mirror, it is useful to construct a *ray diagram*. The ray diagram is based on three rules (you only need two). 1. an incident ray parallel to the principal axis reflects through \\(F\\) 2. an incident ray that passes through \\(F\\) reflects parallel to the principal axis 3. an incident ray which passes through \\(C\\) reflects on itself ## Magnification The calculation of the distance and size of an image are governed by the following equations: $$ \frac{1}{f} = \frac{1}{d_o} + \frac{1}{d_i} $$ where \\(d_o\\) is the distance from the mirror to the object and \\(d_i\\) is the same for the image. The magnification \\(m\\) of an image is the ratio of the image size to the object size: $$ m = \frac{h_i}{h_o} = -\frac{d_i}{d_o} $$ Now, you might be wondering about how the sign of each value affects these equations. The answer is, a LOT. To think about mirrors intuitively, think about what each sign must mean. - \\(f\\) is positive for concave mirrors because the focal point is front, but negative for convex mirrors because it's behind the mirror. - \\(d_o\\) is positive for a real object (in front of the mirror), but negative for a virtual object. A virtual object can't generally exist, but you can simulate it using reflected light. - \\(d_i\\) is the same, but a virtual image makes more sense. - \\(m\\) is obvious, positive for upright and negative for inversion ## Refraction If you put a straw in your water, the straw will strangely seem to bend at the water's edge and go at an angle. The reason for this is **refraction**: light will change direction when it passes through a medium. The change in angle is described by Snell's Law: $$ n_1\sin{θ_1} = n_2\sin{θ_2} $$ where \\(θ_1\\) is the *angle of incidence*, \\(θ_2\\) is the *angle of refraction*, and \\(n_1\\) and \\(n_2\\) are the *indices of refraction* for the corresponding media. The index of refraction is expressed as the ratio between \\(c\\) and the speed of light in the material. If light is very fast in the material, \\(n\\) will be close to 1. When light passes from a medium with a higher index of refraction to a lower one, there's an angle at which the refracted light will be parallel to the media edge, known as the *critical angle* \\(θ_c\\): $$ \sin{θ_c} = \frac{n_2}{n_1} $$ At any angle greater than the critical angle, all the light will be reflected and none will be refracted. > The reason that this is the case can be derived from math. Let's say that light passes from water (1.33) into air (1.00). The critical angle would be equivalent to \\(\sin^{-1}{\frac{1.00}{1.33}}\\) or 48.75°. > > What would happen if we tried to find the angle of refraction based on Snell's Law if the incident angle is 50°, so greater than the critical angle? > > $$ > 1.33 \cdot \sin{50°} = 1.00 \cdot \sin{θ_2} > $$ > > $$ > \sin{θ_2} = 1.33 \cdot \sin{50°} = 1.019 > $$ > > But there's a problem here! The function \\(\sin^{-1}{}\\) is not defined for \\(θ\\) greater than 1, so the angle does not exist. This means that refraction does not occur, the light is only reflected. ## Lens The properties of magnification and refraction are used in lens that are located in glasses, telescopes, microscopes, and even your own eyeball! A **lens** has two mirrors on either side. If the mirrors are convex, the lens is *converging*, and if they are concave, the lens is *diverging*. Every lens has *two* focal points corresponding to each mirror. The ray diagram for a converging lens follows similar rules to a concave mirror, but instead of thinking of reflected rays, the rays are refracted through the lens. 1. When a ray passes through the axis of a lens parallel, it will refract towards the other focal point. 2. When a ray passes through a focal point, it will refract parallel through the other side. 3. When a ray passes through the center of the lens, it will not bend. The rules are the same for a diverging lens, though it may not appear that way. All of the rules for a diverging lens are based on a focal point on the other side of the mirror, but fundamentally have the same property. All of the formulas are the same for lens like mirrors, however, the meaning of the signs changes in this context. - \(f\) is positive for converging lens because the focal point is front, but negative for diverging lens because it's behind the lens - \(d_o\) is positive for a real object (same side as light), but negative for a virtual object. A virtual object can't generally exist, but you can simulate it using reflected light. - \(d_i\) is positive for a real image on the *opposite* side of the light, and negative for a virtual image on the same side of the light. - \(m\) is obvious, positive for upright and negative for inversion ## Lambda Calculus - URL: https://sharifhsn.dev/blog/lambda-calculus/ - Structured data: https://sharifhsn.dev/api/posts/lambda-calculus.json - Description: When we discuss the principles of programming languages, lambda calculus is the bedrock of our discussion. It is the most abstract, formal way to describe functions and application… - Date: 2022-03-03 - Exact published timestamp: 2022-03-03 - Topics: Principles of Programming Languages - Categories: Principles of Programming Languages - Source: Archive - Source URL: None When we discuss the principles of programming languages, **lambda calculus** is the bedrock of our discussion. It is the most abstract, formal way to describe functions and application. ## Turing Completeness A programming language is said to be **Turing complete** if it can compute any function also computable on a Turing machine. It must either be able to emulate a Turing machine, or it must be able to emulate a Turing complete language. Although this is a powerful definition, in truth it is a simple one to satisfy. Features like loops and currying are not related to Turing completeness. ## Syntax Lambda calculus is a very simple language which only uses functions and applications, but is still Turing complete. An expression in lambda calculus is defined as ```bnf e ::= x # variable | λx.e # abstraction (function definition) | e e # application (function call) ``` Everything is a function in lambda calculus. We can make a simple interpreter in OCaml as with Imp: ```ocaml type exp = | Var of string | Lam of string * exp | App of exp * exp ``` in the same order that was defined earlier. `Lam` here is similar to `Fun` for Imp. The lambda calculus expression ```bnf (λx.λy.x y) λx.x x ``` can be deconstructed to its AST as ```ocaml App( Lam("x", Lam("y", App(Var "x", Var "y") ) ), Lam("x", App(Var "x", Var "y") ) ) ``` Let's parse this out forwards. First we read that there is a parenthesis, which means that there is an expression enclosed within. There is a \\(λ\\), which means that the expression is a lambda. The next letter is `x`, which is the name of the `Lam`. The period following separates the function definition to the containing expression, which continues to another \\(λ\\). That means that the containing expression is also a `Lam`, which is called "y". The containing expression is now an `App` which is composed of two `Var`s. This is the end of the nesting because we hit the end parenthesis. The same logic applies for the second expression but with less nesting. Since we have two expressions at the top-level, the full expression is an `App`. The scope of \\(λ\\) extends as far to the right as possible, excepting parentheses. This is why we needed the parentheses for the first term in the above statement, otherwise the first \\(λ\\) would extend throughout the entire statement. The application, however, is left-associative, like OCaml. ## Beta Reduction A function call of type `(λx.e1) e2` replaces all instances of `x` in `e1` with `e2`. That means that we can substitute this statement with `e1{e2/x}`. This is called **beta reduction**. All we have done is apply the function and replace the formal parameters through substitutions. Beta reductions should always be idempotent for the statement. When no more beta reductions can be performed on a term, then it is said to be in *beta normal form*, for example `λx.e`. Another example will be instructive here. Take the term `(λx.λz.x z) y`. This is a function application, since it has two terms. It follows the form that we stated earlier, so we can substitute all instances of `y` on the outside \\(λ\\). This would give us the final term `λz.(y z)` eliminating the outside `λx` and replacing the `x` in the inner term with `y`. ## Alpha Conversion Lambda calculus is **statically scoped** which means that variable definitions are only scoped locally. That means that within a function, you can rename *bound* variables with the same meaning. This is called **alpha conversion**. An important distinction here is between a free and bound variable. Free variables are not contained within a lambda, while bound variables are. For example, in `(λy.λz.y z x)`, `y` and `z` are bound to `λz` and `λy`, respectively, while `x` is free. `λy` contains the `App` of `λz` and `z`. ## More Internet Protocols - URL: https://sharifhsn.dev/blog/other-protocols/ - Structured data: https://sharifhsn.dev/api/posts/other-protocols.json - Description: Besides the primary Internet protocols, there are some others to learn about, like DHCP, NAT, IPv6, etc. - Date: 2022-03-03 - Exact published timestamp: 2022-03-03 - Topics: Internet Technology - Categories: Internet Technology - Source: Archive - Source URL: None Besides the primary Internet protocols, there are some others to learn about, like **DHCP**, **NAT**, **IPv6**, etc. ## DHCP What happens if a device doesn't have a permanent IP address? You take your phone around multiple mobile networks, there isn't a consistent IP address between them. How do you access the internet without an IP address? **Dynamic Host Configuration Protocol** is a client-server protocol that is used in these scenarios. As the name implies, it allows for dynamic IP address allocation that are leased for a certain amount of time. Configuring things like the subnet mask, gateway configuration, etc. to set up an IP is complicated and not feasible for the average user. Imagine if you had to do all that setup everytime your phone moved to a new location! DHCP has two main components: the protocol for delivering the bootstrapping information from the server to clients, and the algorithm for dynamically assigning addresses to new clients. How do you start a connection from nothing? In order to allocate a new address, there are three modes. Automatic allocation gives permanent address, dynamic allocation leases addresses, and manual allocation is managed by a system administrator. DHCP sends packets over a socket at port 67. Its IPv4 header indicates that it is protocol 17. The UDP packet is sent with no initial IP address because there is none. It is sent with the broadcast server IP address `255.255.255.255`. The protocol must also interface with the link layer at the client's MAC address; remember, we're starting at zero. There is also space for the server hostname, boot filename, and several kinds of options. The options indicate what the purpose of the UDP packet is. For example, option 1 is `DHCPDISCOVER`, which is the message that lets the server know that a client is looking for an IP address. The communication sequence is directly encoded into the options in the packet. Some others: - DHCP Offer: server response with parameter proposal - DHCP Request: like discover, but focused to a specific server - DHCP ACK: server gives IP address to client - DHCP NAK: server declines to give IP address to client - DHCP Decline: client declines the given IP address - DHCP Release: client gives up its IP address DHCP typically applies within a subnet. Relay agents on routers, like with BOOTP, allow servers to handle requests from other subnets. ## NAT There are two kinds of IP addresses: *public* and *private*. Private addresses are reserved for `10.0.0.0` to `10.255.255.255`. If you send a request to private address, your router will not send it out to the Internet. Private addresses are also free; you can hand out as many as you want without it costing anything. However, requests sent from a private IP address cannot access the Internet. In order to access the Internet, we use a **Network Address Translation** box. This NAT box is assigned a single public IP address and it is the public-facing IP for all of the machines with private addresses on it. This is likely how your home router works. Each device in your home only has a private address and every time it sends a request to the Internet, that request is sent to the NAT box in your router and translated to one single public IP for all of the machines in your house. The NAT box will take in packets that are sent from private IP addresses and record the packet address in a table, then replace the IP address with its own and its port with some random port. When packets are sent back to that address, the NAT box consults its table and sends the packets back to the correct private IP address based on the port it was sent on. Although there are technically \\( 2 ^ {24} \\) addresses, in practice you can only have less than \\( 2 ^ {16} \\) addresses because each of those needs to have an associated port, and that's a 16-bit number (minus some reserved ports). But why would you do this? Remember that IPv4 is a 32-bit number for addresses, which only allow for \~ 4 billion unique devices. Clearly, there will be, and perhaps already are, more devices that use the Internet than that. IPv6 was created in part to solve that problem, but as many devices only accept IPv4, this works in the meantime. It also provides security to users since ports aren't accessible from the public, they are instead randomly assigned. Even if someone's IP address is tracked, the malicious agent still doesn't necessarily know what device the request is from. ## IPv6 IPv6 is an updated protocol from IPv4 that is optimized for the needs of the modern day. It is significantly simplified compared to IPv4 as much of the machinery in an IPv4 header relates to problems that generally no longer exist. However, the header size is also 40 bytes, twice as large as the IPv4 header. This is because the new source and destination IP address that must be stored in the header are now 16 bytes, not 4, the vast majority of the header size is dominated by address size. IPv6 does not protect against some problems that still exists, such as bit flips. If that happens, the packet is simply dropped because a different layer will do a checksum. | Version | Traffic Class | Traffic Class | Flow Label | Flow Label | Flow Label | Flow Label | Flow Label | | -------------- | -------------- | -------------- | -------------- | ----------- | ----------- | ---------- | ---------- | | Payload Length | Payload Length | Payload Length | Payload Length | Next Header | Next Header | Hop Limit | Hop Limit | followed by the source address and destination address, along 4-bit boundaries ## Electromagnetic Waves - URL: https://sharifhsn.dev/blog/electromagnetic-waves/ - Structured data: https://sharifhsn.dev/api/posts/electromagnetic-waves.json - Description: We have discussed electric fields and magnetic fields, and the way that they change and interact. The full interaction of these fields will create electromagnetic waves, the study … - Date: 2022-03-01 - Exact published timestamp: 2022-03-01 - Topics: Physics - Categories: Physics - Source: Archive - Source URL: None We have discussed electric fields and magnetic fields, and the way that they change and interact. The full interaction of these fields will create **electromagnetic waves**, the study of which will consume the rest of these notes. ## The Electromagnetic Spectrum It's easiest to understand electromagnetic waves (EM waves) in the same way that we understand any other waves, through frequency \\(f\\) or wavelength \\(λ\\). > Throughout these notes, we will refer almost exclusively to the wavelength of an EM wave as identification for consistency's sake. Note however that this is always mapped to a corresponding frequency, though inversely related. Visible light is an EM wave, with wavelengths in the hundreds of nanometers. Stronger EM waves like X-rays or gamma rays go from \\(10^{-8}\\) to \\(10^{-16}\\) meters, which is extremely small! In contrast, weaker EM waves like radio waves can have wavelengths in the thousands of meters. As you can see, EM waves have a broad spectrum. The wavelength of an EM wave in a vacuum is given by $$ fλ = c $$ where \\(c\\) is the well-known *speed of light* at \\(3 × 10^{8}\\) meters per second. \\(c\\) can also be more precisely defined as $$ c = \frac{1}{\sqrt{ε_0μ_0}} $$ where \\(ε_0\\) is the permittivity of free space \\(8.85×10^{-12}\\) and \\(μ_0\\) is the permeability of free space \\(4π×10^{-7}\\). ## Energy EM waves transport energy, and that energy has a certain value depending on the electric and magnetic fields that created the EM wave. The *electric energy density* is: $$ u_E=\frac{1}{2}ε_0E^2 $$ In contrast, the *magnetic energy density* is given by: $$ μ_B=\frac{1}{2μ_0}B^2 $$ Notice the different units here. The electric energy density is dependent on permittivity, while the magnetic energy density is dependent on the inverse permeability. We can combine these two to get the total energy density: $$ u = \frac{1}{2}ε_0E^2 + \frac{1}{2μ_0}B^2 = ε_0E^2 = \frac{B^2}{μ_0} $$ The direction of the wave must always be mutually perpendicular to the direction to the electric field and the magnetic field. This follows the **right hand rule** where the thumb is the magnetic field \\(B\\), the pointer finger is the propagation of the wave, and the middle finger is the electric field \\(E\\). You can remember this because the pointer finger points to where the wave is going, and \\(B\\) comes before \\(E\\) in the alphabet, so the first one is the thumb and the second is the middle finger. > An EM wave has a magnetic field with an rms value of \\(3.40 × 10^{-6} T\\). The wave passes perpendicularly through an opening that has an area of \\(0.35 m^2.\\) > > To get the electric field value \\(E\\), we multiply the magnetic field \\(B\\) by \\(c\\) to get \\(3 × 10^8\ m/s\ •\ 3.4 × 10^{-6} T = 1020\ N/C\\). > > To get the energy densities \\(μ_E\\) and \\(μ_B\\), we simply apply the earlier formulas here. > > $$ > μ_E = \frac{1}{2}ε_0E^2 = \frac{1}{2}(8.85×10^{-12})(1020)^2 = 4.604 × 10^{-6} J/m^3 > $$ > > $$ > μ_B = \frac{B^2}{2μ_0} = \frac{(3.4×10^{-6})^2}{2\ •\ 4π×10^{-7}} = 4.6×10^{-6} J/m^3 > $$ > > After getting the electric energy density and the magnetic energy density, getting the total energy density is as simple as adding them together to get \\(9.204×10^{-6}\ J/m^3\). > > The intensity of a wave is a simple equation; you can think of it as the density multiplied by speed being how much is being transferred per unit time: \\(S = cu\\). > > $$ > S = cu = (3×10^8 m/s)(9.204×10^{-7} J/m^3) = 2761.2 W/m^2 > $$ > > Let's say we want to find the energy that is carried through this opening over twenty seconds. Energy is in the unit joules \\(J\\) and \\(W\\) is \\(J/s\\). In order to convert a dimensions properly, we're going to need multiply the intensity by \\(s\ •\ m^2\\). With this in mind, it's clear how we should form our equation. We're multiplying the density by the time passed (20 seconds) and the area \\(0.35 m^2\\). > > $$ > E = StA = (2761.2 W/m^2)(20 s)(0.35 m^2) = 19328.4\ J > $$ ## Polarization Electromagnetic waves can be **polarized** in a a particular direction by passing throw a material which only allows for one vector to pass through. Unpolarized light passing through a polarizing material will have *half* the intensity that it originally had. **Malus' Law** states that an *analyzer* which alters the intensity and polarization direction of an EM wave will decrease the intensity as so: $$ \bar{S} = \bar{S_0}\cos^2{θ} $$ and changing the polarization direction to match the analyzer. \\(θ\\) in this case is the *difference* between the polarized light and the analyzer. Say the light is polarized at angle of \\(30°\\) clockwise to the vertical and it passes through a filter that is at an angle of \\(15°\\) counterclockwise to the vertical. The \\(θ\\) in this case will be \\(45°\\). > A vertically polarized beam of intensity \\(S_0 = 60.0\\) is incident through three polarizers \\(θ_1 = 38.0°\\) counter-clockwise, \\(θ_2 = 17.0°\\) clockwise, and \\(θ_3 = 30.0°\\) counter-clockwise. > > $$ > S_1 = S_0\cos^2{θ_1} = 60\cos^2{38°} = 37.258\ W/m^2 > $$ > > $$ > S_2 = S_1\cos^2{θ_2} = 37.258\cos^2{55°} = 12.257\ W/m^2 > $$ > > $$ > S_3 = S_2\cos^2{θ_3} = 12.257\cos^2{47°} = 5.701\ W/m^2 > $$ ## ISP Addressing - URL: https://sharifhsn.dev/blog/isp-addressing/ - Structured data: https://sharifhsn.dev/api/posts/isp-addressing.json - Description: How does an ISP assign IP addresses to its customers? - Date: 2022-02-28 - Exact published timestamp: 2022-02-28 - Topics: Internet Technology - Categories: Internet Technology - Source: Archive - Source URL: None How does an ISP assign IP addresses to its customers? An ISP has a block of addresses that are partitioned to its customers. If an ISP network has an address like `200.8.4/24` address, that is 256 addresses. `/20` is for 4K hosts, and `/16` is for 64K hosts. In fact, the calculation is \\( 2^{32 - n} \\). ## Subnetting A network can be subdivided into **subnets**. This way you can have each router handling a smaller portion of the network, or have different kinds of routers i.e. wired/wireless handling different subnets. In order to divide IP addresses, we use a **subnet mask**. For example, if our network is `128.64.32/24`, a range we could have is `128.64.32.0-127` and `128-255`. The easiest way to mask this is to look at the most significant bit. In the first range, the most significant bit is 0, and in the second range, the most significant bit is 1. The mask in this case would be `255.255.255.128` or in hex, `F.F.F.8`. If we bitwise AND this mask with the IP address, we will get the correct subnet. We can follow the same procedure to get further divisions by the power of 2. We can mask over two bits to get four subnets of ranges `0-63`, `64-127`, `128-191`, and `192-255` with a mask of `255.255.255.192`. Let's think of an analogy. If we want to deliver some mail to our neighbor, the easiest way to do it is to go directly to his mailbox and give it to him instead of passing it off to the post office. In the same way, we can quickly send messages between hosts on the same subnet. In order to facilitate this, the router will first check if the sender and destination (both first &ed with the subnet mask) are in the same subnet. If they are, it doesn't bother sending the message over the Internet and it will actually just directly send the message to the destination itself. In particular, this makes email between people on the same networks extremely quick. This is why it's nice to have email between people with the `scarletmail.rutgers.edu` domain. ## IPv4 Header In order to get a packet to a destination host, we need a packet header with the identity of both the destination identity and the source identity. At this layer, we don't worry about reliable exchange, dropping packets, sequence numbers, etc. That's for TCP to handle, our only job is to send the packet to the destination. However, we have to detect some kinds of problems. We don't want to loop packets over and over, this will overload the network. There used to be a worry of fragmentation, if the size of the packet is greater than the *maximum transmission unit* (MTU) of the router. Nowadays, everyone has broadband so this is not an issue because everybody has high link speeds. When IPv4 was invented we needed checksums for verification if there are bit flips. IPv6 addresses have much less checksum machinery because it's a waste nowadays, it only checks at the end. IPv4 has to calculate a checksum at every router, which catches errors early but is generally wasteful. The *time to live* (TTL) is 1 byte long and is used to protect against loops. It starts at 255 and decrements every time it passes through a router. If the TTL is 0, the packet is dropped and a "time exceeded" error is thrown to sender. This prevents the packet from being looped around forever. An entire 4 bytes is devoted to fragmentation machinery. This is not in IPv6, which simply throw an error if the packet is too large. The reason this is fine because it is a rare case nowadays and it cuts fat out of the header. Here are the steps of fragmentation: 1. router receives packet larger than MTU 2. if *Don't Fragment* (DF) flag is set, throw *Fragmentation Needed* error 3. else, divide the packet into fragments maximum size `outgoing MTU` - header size (20 bytes) 4. in each new packet fragment, the `length` field is the size of the fragment, the *More Fragments* (MF) flag is set except for the last fragment, and the `fragment offset` field is set to the offset of the fragment in 8-byte blocks. The checksum is recomputed for the fragment. [IPv4 - Wikipedia](https://en.wikipedia.org/wiki/IPv4#Fragmentation_and_reassembly) ## ICMP What do we do to propagate errors? We're already using IP for sending messages, how do we send errors? This what the **ICMP** protcol is for. It is unreliable like UDP, and is tightly coupled with the implementation of IP. Its protcol ID is 1, which signifies its importance. There are certain known error codes. The *echo request/reply* also known as *ping* is simply a request to check if the host is alive and responding. There is also the **traceroute**, which records the route that a packet takes by tracking TTL. As we know, when a router receives a packet, it decrements TTL. If we start TTL at 0 and slowly increase it, it will throw time exceeded back to us from each router in order. This way, we can see every router on the way to our destination. ## Making a Website with Zola, Github Pages, and Github Actions - URL: https://sharifhsn.dev/blog/making-a-website/ - Structured data: https://sharifhsn.dev/api/posts/making-a-website.json - Description: Making a website in the modern era is not easy to do for free. I did it using Zola, Github Pages, and Github Actions. I've always wanted to have a personal website where I can uplo… - Date: 2022-02-26 - Exact published timestamp: 2022-02-26 - Topics: zola, github actions, website, Meta - Categories: Meta - Source: Archive - Source URL: None Making a website in the modern era is not easy to do for free. I did it using Zola, Github Pages, and Github Actions. I've always wanted to have a personal website where I can upload what I do on my local computer to access remotely and have the world see. But it always seemed like too much of a hassle to set up and I didn't have the capital to invest in a website that I didn't need. However, when I started going back to university in-person this year, I found it much easier to take notes by typing them instead of using OneNote as I was accustomed, as the amount of code I had to write was drastically increased. Needing a way to access them remotely with a nice view, I thought a blog would be a good way to do that in addition to all the other things I had always wanted a website for. So, I embarked on a journey to create a website. ## Github Pages A website is no use unless we have somewhere to put it. [Github Pages](https://pages.github.com/) is a service offered by Github since 2008 that allows you to host your own website from a Github repository. You get one free website per Github account, which is called [username].github.io. All we have to do to enable it is create a repository named [username].github.io and enable Github Pages in the settings! ```bash # should be above 2.28 to enable default branch name change git --version mkdir [username].github.io cd [username].github.io # personal git config git config --global user.name "NAME" git config --global user.email "EMAIL" git config --global init.defaultBranch "main" git init gh repo create [username].github.io --public --source=. --remote-upstream ``` If you're using [Visual Studio Code](https://code.visualstudio.com/) as your editor, there's a nicer way to do this than through the command line. After installing [the Github extension](https://marketplace.visualstudio.com/items?itemName=GitHub.vscode-pull-request-github), go to the Source Control button on the sidebar. There should be a button labeled "Publish to Github" which allows you to interactively initialize a Git repository in the current folder and publish it to Github. ## Github Actions We might have created our Github page, but we need a way to get all of the code from our repository to the website. This is called **deployment**. Luckily, we have a way to automatically deploy our website through **Github Actions**. [Github Actions](https://github.com/features/actions) is another service offered by Github since 2019 that gives you free CI in public repositories. Although the free tier is [fairly limited](https://docs.github.com/en/billing/managing-billing-for-github-actions/about-billing-for-github-actions#included-storage-and-minutes) at 500 MB and 2000 minutes per month, it should be more than enough for a static blog that is not deployed very often. There is an action automatically created for our Github page called `pages-build-deployment` which, as the name implies, builds and deploys the page you've created on push. The way that I organized my code, which is probably the simplest option, is that I hosted my code at the `main` branch and had a `gh-pages` branch that hosted the actual website which was built from the `main` branch. If you want to do the same, go to `Settings` -> `Pages` and make sure that the build target is the `gh-pages` branch at the root. For now, this won't do anything because we don't have a `gh-pages` branch or anything in our `main` branch. So how do we *actually* make our website? ## Zola [Zola](https://www.getzola.org/) is a static site generator written in [Rust](https://www.rust-lang.org/), and is one of the fastest out there. I decided to choose it for my website. If you'd prefer a different generator, this is where this guide diverges for you. There are plenty of tutorials for Hugo websites or others, but I have found a lack of Zola guides so I decided to create this. To start, [install Zola on your system](https://www.getzola.org/documentation/getting-started/installation/). The documentation on the website is pretty stellar so I would recommend reading that to get a quick understanding on how to use Zola. I will explain the parts that I personally found difficult or unclear. ```bash zola init zola build # unnecessary as serve will also automatically build it zola serve ``` Now you can see your new website at `127.0.0.1:1111`! ## Zola Deployment Although we can build your website very simply on our local machine, it would be preferable to automatically build the website when we publish content so we don't have to mess around with all of that. The [Zola-approved way](https://www.getzola.org/documentation/deployment/github-pages/) to do this is by using [zola-deploy-action](https://github.com/shalzz/zola-deploy-action). All of you have to do is click the `New Workflow` button on the `Actions` page from your Github repository and follow the link to `set up a workflow yourself`, then copy-paste this into it: ```yaml # On every push this script is executed on: push name: Build and deploy GH Pages jobs: build: runs-on: ubuntu-latest if: github.ref == 'refs/heads/main' steps: - name: checkout uses: actions/checkout@v2 - name: build_and_deploy uses: shalzz/zola-deploy-action@v0.14.1 env: # Target branch PAGES_BRANCH: gh-pages # Provide personal access token TOKEN: ${{ secrets.TOKEN }} ``` However, I wanted more configuration and control over my action. Github Actions operates using Jekyll by default, so unless you add a `.nojekyll` file to the build branch it will run unnecessary steps to build a Jekyll theme. In order to reduce complexity, I decided to make a similar action that adds that command. If you're comfortable with the workflow as provided, then skip the next section. ## Creating a Github Action An action of the type we want here consists of three files: `action.yaml`, `Dockerfile`, and `entrypoint.sh`. Let's break these down. - `action.yaml`: the configuration file that tells Github Actions what to do - `Dockerfile`: the configuration file for the Docker container that Github Actions will set up - `entrypoint.sh`: the shell script that will execute the commands we want `action.yaml` is simple. It should follow this general format: ```yaml # action.yaml name: 'ACTION_NAME' description: 'DESC' author: 'NAME' runs: using: 'docker' image: 'Dockerfile' ``` The Dockerfile is more complex and has many more options. I kept mine simple to what is needed, you may have your own preference for Docker images. ```Dockerfile # any Docker image is fine, I prefer debian from debian:stable-slim MAINTAINER NAME # for github actions LABEL "com.github.actions.name"="ACTION_NAME" LABEL "com.github.actions.description"="DESC" # locale, I am in the U.S. so I use en_US ENV LC_ALL C.UTF-8 ENV LANG en_US.UTF-8 ENV LANGUAGE en_US.UTF-8 # standard apt-get + wget and git for getting and building RUN apt-get update && apt-get install -y wget git # get zola on the docker image RUN wget -q -O - \ "https://github.com/getzola/zola/releases/download/v0.15.3/zola-v0.15.3-x86_64-unknown-linux-gnu.tar.gz" \ | tar xzf - -C /usr/local/bin COPY entrypoint.sh /entrypoint.sh # give the entrypoint executable permissions RUN chmod +x entrypoint.sh ENTRYPOINT ["/entrypoint.sh"] ``` `entrypoint.sh` is where all the magic happens. Fundamentally, all that it does is call `zola build` on the `main` branch, which builds your website inside the `public` directory (you can configure this if you so wish). It then commits those website files to the `gh-pages` branch. We need one more thing before we can create `entrypoint.sh`; a token. We need to authorize our action to be able to push to `gh-pages`. You can create a token by going to [this page](https://github.com/settings/tokens) and creating a new token with at least `repo` rights. You can add the token to your repository by going to `Settings` -> `Secrets` -> `Actions` and creating a new repository secret called `TOKEN` (or any other name you like). With that token, this is the basic necessities for `entrypoint.sh`: ```bash #!/bin/bash set -e set -o pipefail main() { git config --global url."https://".insteadOf git:// git config --global url."$GITHUB_SERVER_URL/".insteadOf "git@github.com": # update git submodules (important if you have themes) git submodule update --init --recursive zola build cd public # if you want to add any commands do it here e.g. `touch .nojekyll` git init git config user.name "GitHub Actions" git config user.email "github-actions-bot@users.noreply.github.com" git add . git commit -m "Deploy ${GITHUB_REPOSITORY} to ${GITHUB_REPOSITORY}:gh-pages" git push --force "https://${GITHUB_ACTOR}:${TOKEN}@github.com/${GITHUB_REPOSITORY}.git" master:gh-pages } main "$@" ``` With all of these settings, this is how your workflow should look: ```yaml # .github/workflows/main.yml on: push name: Build and deploy GH Pages jobs: build: runs-on: ubuntu-latest if: github.ref == 'refs/heads/main' steps: - name: checkout uses: actions/checkout@v2 - name: build-and-deploy uses: ./ # wherever your action is in relation to the root of the repo env: TOKEN: ${{secrets.TOKEN}} ``` ## Themes Zola requires a [Tera](https://tera.netlify.app/) template to render your site for the base site `index.html` as well as `page.html` for page-specific settings. You can also use [Sass](https://sass-lang.com/) stylesheets if you enable it in your `config.toml`. **Themes** are a convenient way to have those built for you so you can get a website looking nice without excessive fiddling. I decided to use the [after-dark](https://github.com/getzola/after-dark) theme which is based on the Hugo theme of the same name. Since I wanted to make my own modifications to it, [I forked it](https://github.com/sharifhsn/after-dark) and added the changes I wanted. The easiest way to add a theme for Github pages is to use submodules: ```bash git submodule add https://github.com/getzola/after-dark.git themes/after-dark ``` The deploy action will take care of updating the submodule as necessary. Just add the theme to your `config.toml` file and voila! Although I like the `after-dark` theme, I might change to a different theme or make my own in the future to accommodate my goals for this website. If that happens, I'll detail that process in another post. ## $\KaTeX$ Most of the lecture notes I write incorporate [$\KaTeX$](https://katex.org/) in some way. I find it an expressive way to write formulas and math expressions when reviewing for tests. `after-dark` does not provide $\KaTeX$ support by default, which is part of the reason I forked it. My preferred option for $\KaTeX$ rendering would be server-side, as I don't plan on pushing very often (perhaps once per day) and Zola compilation is extremely quick. However, after doing some research into [previous attempts](https://github.com/getzola/zola/pull/1073), I decided it wasn't feasible for now. Perhaps in the future I'll take a stab at implementing it myself, but for now I'll settle for client-side. [The $\KaTeX$ docs](https://katex.org/docs/browser.html) give a pretty good description on how to incorporate it into your website. In Tera, all you have to do is enclose those stylesheets/scripts into CSS/JS blocks, respectively. My inspiration came from [this pull request](https://github.com/getzola/after-dark/pull/22). I modified the standard `auto-render.min.js` script to add standard $\KaTeX$ \$ \$ tags. ## Currying Arguments in a Function - URL: https://sharifhsn.dev/blog/curry/ - Structured data: https://sharifhsn.dev/api/posts/curry.json - Description: Currying is the concept of having multiple arguments in a function. OCaml defaults to currying its functions. int - int - int is a function that takes two ints and returns an int. - Date: 2022-02-21 - Exact published timestamp: 2022-02-21 - Topics: Principles of Programming Languages - Categories: Principles of Programming Languages - Source: Archive - Source URL: None **Currying** is the concept of having multiple arguments in a function. OCaml defaults to currying its functions. `int -> int -> int` is a function that takes two `int`s and returns an `int`. The `->` is *right-associated* and the function application is *left-associated*. The last element of the function definition is always the return type, but calling a function always counts arguments from the left. ```ocaml let f a b = a / b;; let f = fun a -> (fun b -> a / b);; ``` These two lines are equivalent because the second line is just the uncurried function separated into two parts. The `fun` lambda only has one argument, and the `->` keyword is right-associated so `b` is considered part of the arguments. Currying allows you to pass only a portion of the expected arguments to the function, the same way that Python uses keyword arguments. Another way to enable multiple arguments is by using a tuple that contains both arguments, which are destructured in the function definition. However, the advantage of currying is that you can separate the call of the function from the arguments. ```ocaml let add a b = a + b;; let addthree = add 3;; addthree 4;; (* evaluates to 7 *) ``` This code allows `addthree` to exist as an implementation of `add` with a specific argument already given. However, it's not all roses with currying. Function need to retain state regardless of stack state, so the local variable `3` that is temporarily in the function `addthree` may not always be there. Anonymous functions may not have the same call stack. In C-like languages, local variables are contained within their own stack frames. When a function calls another function, the new stack frame that is created contains its own local variables. If a variable is declared, then initialized through a function call, that variable contains junk until the function returns. If we return a function that references a local variable in another function, reading it off the stack can get confusing. The first variable you look for is still uninitialized, so you need to evaluate where that variable is coming from to understand. We solve this problem using **static scoping**. Nonlocal names refer to their nearest binding in the program text. This is also known as lexical scoping. If two variables have the same name in an inner scope and an outer scope, then we read the one in the inner scope first. ## HTTP Explained - URL: https://sharifhsn.dev/blog/http/ - Structured data: https://sharifhsn.dev/api/posts/http.json - Description: HTTP is probably the most visible Internet protocol to end users. It appears (as well as its cousin, HTTPS) at the beginning of every URL we use to access the internet. But how doe… - Date: 2022-02-20 - Exact published timestamp: 2022-02-20 - Topics: Internet Technology - Categories: Internet Technology - Source: Archive - Source URL: None **HTTP** is probably the most visible Internet protocol to end users. It appears (as well as its cousin, **HTTPS**) at the beginning of every URL we use to access the internet. But how does it work? HTTP stands for HyperText Transfer Protocol. It defines the structure of messages that are passed between two programs: a *client* and a *server*. But before we talk about HTTP, let's clarify some vocabulary ## Web Terminology - **object** - a file that has an associated URL - **Web page** - a document that consists of multiple objects, typically a base HTML file and other objects. The HTML file will reference other objects through a path. - **URL** - address consisting of the following parts: | protocol | hostname | path name | |:--------:|:------------------:|:---------------------------:| | http:// | www.someSchool.edu | /someDepartment/picture.gif | - **Web browser** - an application such as Firefox or Chrome that acts as the client in HTTP - **Web server** - an application such as Apache that acts as the server in HTTP and houses the Web objects in question ## HTTP Low-Level Under the hood, HTTP uses TCP to make a reliable connection between the server and client. The client and server both send requests and receive responses from their respective socket interfaces. When a file is sent, there is no information that the server inherently remembers about the HTTP connection. Because of this, HTTP is considered a **stateless protocol**. This is not always desirable. Sometimes, websites want to remember clients in order to ease use of a website. We will discuss workarounds later. ## Connections The original HTTP 1.0 protocol was *non-persistent*. This means that every request made between a client and a server had its own connection that was opened and closed every time a request was made. Here's the process: 1. A TCP connection is created between the client and server on port 80. 2. An HTTP request message is sent from the client to the server through the socket. 3. The server receives the message, retrieves the object that it is requesting, and sends the response message to the client through the socket. 4. The server "closes" the TCP connection (TCP will wait until it knows that the client has received the message) 5. The client receives the message and the connection terminates. It will extract the HTML file from the message and initiate a new request for each object referenced within, repeating these steps as needed. Clearly, this method is inefficient. We are opening a TCP connection for every single request that is made between the same client and server. This inefficiency is mediated slightly by the ability to open serial TCP connections; typical web browsers open 5 to 10 parallel TCP connections. The response time for each connection is \\(2 \\cdot RTT + T_{trans}\\). HTTP 1.1 introduced *persistent* connections. In step 4, the server doesn't close the TCP connection after sending the response. The entire Web page is sent over the same TCP connection, and the requests can be pipelined for further performance gains. The connection is typically closed on timeout. HTTP/2 further builds on this by interleaving requests and responses in the same connection with a way to preserve ordering on each end. ## Message Format ```http GET /http HTTP/1.1 Host: sharifhsn.github.io Connection: close User-agent: Mozilla/5.0 Accept-language: fr ``` The `Host` line is technically unnecessary, since the connection is already established by this point, but it's useful for Web proxy caches. The `Connection` line means this connection should be non-persistent. The `User-agent` line specifies the **user agent**, which is the browser type that is making the request to the server. This is used so that the server can send different websites to different browsers. `Accept-language` expresses a preference for content in French if it exists, otherwise give the default. In general, the header lines exist to give information that is necessary for the request. The HTTP response is different than the request. An example response to this message might be: ```http HTTP/1.1 200 OK Connection: close Date: Tue, 09 Aug 2011 15:44:04 GMT Server: Apache/2.2.3 (CentOS) Last-Modified: Tue, 09 Aug 2011 15:11:03 GMT Content-Length: 6821 Content-Type: text/html ``` The status line at the top tells us some similar information, like the protocol we are using. The `200 OK` is our status code for a good response. There are other status codes: - 200 OK: request succeeded, information returned. - 301 Moved Permanently: the address is moved, for example when a website links to its www address; there will be a location header telling the client where the new place is to go - 400 Bad Request: request is written poorly and could not be understood - 505 HTTP Version Not Supported: self-explanatory lol There's tons of valid header lines we can use. The HTTP specification allows for many header lines to be used from various software. Writing a correct HTTP implementation isn't easy. ## Cookies I mentioned before that HTTP is stateless, which means that the server does not preserve any information about the client after the connection ends. However, websites really want ways to identify users for operations such as logins. In order to facilitate state, websites use **cookies**. You might have heard of them when you log in to a website and it asks you to accept cookies. The typical creation method of a cookie might look like this: 1. A client makes a request to the server for the first time. 2. The server generates an identification number for the client. 3. The server sends a HTTP response to the client with a cookie in the header: `Set-cookie: 1678` 4. The client browser stores the cookie in a special file and sends a request to the server with the same cookie number. 5. Now, whenever the client requests the server, it will include a cookie in the header which the server will use. ## Caching You might have heard of a **proxy server** before. These servers are usually used by ISPs like your home ISP and your university ISP. These servers contain frequently-accessed data from several origin servers and are located much more closely to the client. These servers are extremely helpful in improving response times for the client as well as reducing traffic for the server. The close location and LAN nature of proxy servers mean that the request/response time between clients and servers are quick. And since the origin server is not accessed as often, its burden of traffic is significantly reduced. The delay in fetching a response from the Internet tends to be dominated by the time for the institutional server receiving a response from the Internet. However, if there is heavy traffic on the *access link* between the institutional server and the Internet router, that can slow down speeds to minutes per request, which is obviously unacceptable. $$ requestRate \\cdot requestSize / speed = trafficIntensity $$ The intensity between client and LAN is usually negligible, but the access link between LAN and Internet dominates. With a Web cache, hit rate is generally between 0.2 and 0.7. $$ cacheHit \cdot LANspeed + (1 - cacheHit) \cdot internetSpeed$$ The traffic intensity stops being an issue and Internet delay dominates again once a cache is introduced. It's less expensive and faster than increasing access link speed. ## Paging - URL: https://sharifhsn.dev/blog/paging/ - Structured data: https://sharifhsn.dev/api/posts/paging.json - Description: The main problem that segments have introduced to managing memory space is that their variable size wastes memory through fragmentation. Fixed-size pieces that are easier to handle… - Date: 2022-02-20 - Exact published timestamp: 2022-02-20 - Topics: Operating Systems Design - Categories: Operating Systems Design - Source: Archive - Source URL: None The main problem that segments have introduced to managing memory space is that their variable size wastes memory through fragmentation. Fixed-size pieces that are easier to handle are much more popular: these are known as **pages**. ## Pages Address spaces are split up into multiple pages, typically a power of 2. For example, let's picture a tiny 6-bit address space that can only address 64 bytes. We could split it up into four 16-byte pages, that we'll refer to in sequence. The physical representation of these pages is through **page frames**, which are sequenced directly in memory and are ordered. The pages in the virtual address space map directly to page frames in physical memory, with no respect for order. Memory management of the free space here can be done with a simple free list. All it needs to look for is four free page frames *somewhere* in memory and map each page to a page frame. This mapping is stored in a **page table** which is kept *per process*, since every process has its own address space. ## Page Translation (IMPORTANT) In our example from earlier, we worked with a 6-bit address space. The highest order bits are reserved for the virtual page number, and the lower bits are reserved for the offset within the page. For example, the virtual address 21 would be `010101` in binary. The top two bits `01` tell us that we are looking for virtual page 1. The OS checks the page table and sees that VP2 maps to physical frame 7, which is `111` in binary. The OS then translates the physical address by replacing the VP bits with the PF bits. In this example, the new address would be `1110101`, or the 117th byte in memory. ## Page Tables We've talked about page tables a bit, but let's go into details. They are a data structure like any other, but they can get very large. Each mapping in the table, which is called a **page table entry (PTE)**, is typically \\(2^2\\) bytes in size. It's often simpler to think of these calculations in terms of the bits involved, since they will always be a power of two. With that in mind, this is the formula for page table size: $$ pageTableSize = VPN + PTE $$ That's 22 bits for a typical 32-bit system, which is pretty massive at 4MB per page table, which is again per process. We obviously can't keep this in the MMU, so we need to actually store it in memory. How do we organize the page table as a data structure? The most obvious way is as a linear table, aka an array. The VPN is an index in the array, and the value at the index is the PTE which gets the PFN. The PTE itself contains the PFN plus some helpful bits for us. There is a valid bit checks if the mapping even exists, protection bits for privileged memory, and others. This access is still extremely slow for load store operations. We need a better solution! ## TLB The hardware comes to the rescue here. MMUs provide a **translation-lookaside buffer** to cache commonly used translation. When a load-store instruction is executed, the CPU will first check the TLB if there exists a translation for the specified address if it exists. If it does, then it can quickly perform the operation without performance overhead of going back and forth on memory so often. If it doesn't, the CPU hands the reins over to the OS to do its own page and offset translation and it will cache the translation after the OS gives it. Here's the basic steps: - extract VPN from virtual address - check if VPN is in TLB - if hit, then perform the PFN concatenation and access the memory - if miss, hardware will check the page table to find the translation - update the TLB with the translation - retry TLB, now a guaranteed hit All of this assumes that memory is valid, accessible, and unprivileged; these can cause exceptions that the OS will handle. Step 4 is expensive and it is therefore the step that we want to avoid as much as possible. Luckily, many memory accesses reside on the same page. For example, arrays are almost guaranteed to have the same TLB hit because they are organized contiguously. **Spatial locality** helps us increase our hit rate. Page size is also significant here. By having big pages, typically 4KB, we are unlikely to have TLB misses since the same page is accessed for this memory. Temporal locality will also help here, as the TLB will evict based on recency. The TLB miss can be managed by either the CPU or the OS. Older **CISC** or complex-instruction set computers managed the TLB themselves, while modern **RISC** reduced computers raise an exception to be handled by the OS. In these computers, steps 4 and 5 are handled by the OS and step 6 only executes after the exception is handled. The exception raised here is a little different from other exceptions. Typically, instructions that raise exceptions are skipped. Here, we want to retry the operation, so we need to get a different program counter. The TLB cache is **fully associative** which means that any translation can be anywhere. This means that the VPN is encoded with the PFN and other bits. One small note is that the TLB entries have valid bits just like the PTE, but they serve different purposes. In the TLB, it refers to a valid translation. If a context switch has occurred, for example, then all of the cache becomes invalidated. PTE valid bits refer to unallocated memory which results in a process kill. Let's examine that cache invalidation. Flushing the cache on every context switch seems to miss the point of a TLB since it happens so often. What are other ways we can manage this? Hardware will typically add an **ASID** (address space identifier), which is similar to a PID but contains less information. This way, a TLB can contain information for multiple processes without accidental contamination. ## Smaller Tables As discussed earlier, page tables can get real honking big, which is not good for memory consumption. TLBs can mitigate the performance problems of page table *access*, but we need better ways to mitigate memory usage of our page tables. The obvious solution, also mentioned earlier, is to just increase the size of each page. The size of a page in \\(2^{bits}\\) is given by: $$ addressSpace = numberOfPages + pageSize $$ so if we have a 32-bit address space, we can have 18 bits for our number of pages and 14 bits for our 4 KB page. 1 MB per page table is better because it's smaller, but there's a limit to this kind of strategy. If we make our pages too big, we will get **internal fragmentation** within each page, where an entire page is allocated but not that much memory within it is used, leading to waste. 4 KB is a good middle ground, which is why that's what x86 uses. ## Segmentation-Paging We looked at segmentation earlier, but initially dismissed it due to its issues of variable size. What if we combined the two approaches into a hybrid? Instead of having one giant page table for the address space, what if we split it up into segments? We can twist around the base/bounds registers to use them for a different purpose. The base register can hold the physical address of the actual page table for a particular segment, and the bounds will tell us where the physical end of the page table is. We can use these values to calculate the number of valid pages. Again, let's use the top two bits to refer to segment. We'll figure out our base/bounds pair from the those bits, then get to our PTE by adding \\(VPN \cdot sizeof(PTE) \\) to it. The bounds register can be used to track our number of valid pages so they don't take up space in the page table if they are not used. However, this comes with the same issues of segmentation earlier: external fragmentation. ## Multi-level Page Tables As is common for many problems, the solution for page tables is to put them inside another page table. **Multi-level page tables** are the de facto solution for page tables that are used in x86. However, this introduces a significant amount of complexity in our page table search, so we will need to examine that. We will chop up our page table into its own kind of pages, each of which only contain PTEs. If an entire page of PTEs is all invalid, which means that none of them have valid translations, don't allocate that page. In order to manage this apparatus, we will introduce a new data structure: the **page directory**. You can think of the page directory as a simple page table if the only memory being tracked was the sub-page table. Each entry has a valid bit which is true if *any* PTE in its mapped page is valid, as well as the PFN where the page of PTEs is stored. That PFN is only allocated if the valid bit is set. If we organize this structure correctly, each chunk of the page table that we refer to as "pages" are actually page-sized and can fit into memory pages in kernel space. This greatly simplifies page table management. However, there is a cost, as always. TLB misses require two memory loads in order to get the correct page because of the level of indirection that page directories introduce. Since TLB misses are rare, we take that tradeoff in return for significantly reduced memory consumption. This is a trade for space that sacrifices time. Let's get some numbers for a multi-level page table. Assume a 14-bit virtual address space split into 8 bits for VPN and 6 bits for offset. Remember that this translates to 256 entries per table and a page size of 64 bytes. We have to do some special bit magic to manage these page levels as well. In this example, our page table size is 10 bits; remember, it's VPN + PTE. We need to subtract our offset bits from that size to get 4 bits for number of pages, then subtract our PTE bits from that to get number of PTEs per page. In total, here is the formula: $$ pageTableSize = numberOfPages + \underbrace{pageSize}_{PTESize + PTEPerPage} $$ \\(10 = 4 + (2 + 4)\\) in this example. The page directory size is the same as the number of pages, so it is also 4 bits. When doing our translation, the highest order bits of the VPN are reserved for the page directory index. After indexing using these higher bits, the rest of the bits of the VPN are the "offset" within the page that the higher bits point to. You can think of this is as a mini address with the VPN being the PDI and the offset being the page table index (PTI). ## Infinity and Beyond This example presented has only two levels, the page directory and the page table. But we can go deeper. Let's reset our numbers to 30 bit address space and 9 bit page size. The VPN is therefore 21 bits. If we apply the same formula as before here, we end up with 7 bits for PTEPerPage. The lowest bits of the VPN will be reserved here. Our page directory will now have 16 bits (14 + 2 for each entry), but this is WAY too much. We need every piece of our structure to fit into a page, so we need to get the PDI to 7 bits or less. The solution here is to further split the PDI into another VPN/offset split. The offset here will be 7 bits, as this is what fits into the page. The zeroeth index is the VPN part, which is 7 bits. It can address 128 pages of the second-level directory. The second-level directory will address 128 pages of PTEs. This way, both the first and second level directory are 9 bits (index + entry size) so they fit into a page! This is *really hard to understand!* Try to think of it where every level is a split between index and offset. The first offset must always be the size of a page, and every inner offset must be the size of page minus the PTE size. In this case, the first offset was 9 bits, and every directory offset was 7 bits. To get the maximum level \\(n\\), the formula is: $$ addressSpace > (pageSize - PTESize) \cdot (n - 1) + pageSize $$ Our maximum level in this example is 3, because a level of 4 would result in \\((9 - 2) \cdot (4 - 1) + 9\\) which is 30, the same as the address space, which must be greater. ## Inverted Page Tables A small coda to this discussion of paging is the **inverted page table**. Unlike other kinds of page tables, this is a page table that is shared between processes that maps *every physical page*. This massive table has information per entry about what process is using it, the virtual page number that the process using is referring to, and the physical page. PowerPC uses this model, with a hash table instead of an array to speed up lookups. However, lookups are still slow, although they take up less memory. It is also difficult to implement sharing as specified earlier. You need to somehow chain multiple virtual addresses to one entry, which introduces serious complexity. ## TCP Explained - URL: https://sharifhsn.dev/blog/tcp/ - Structured data: https://sharifhsn.dev/api/posts/tcp.json - Description: TCP is the well-known protocol for reliable transfer of data. But how does it actually work? - Date: 2022-02-20 - Exact published timestamp: 2022-02-20 - Topics: Internet Technology - Categories: Internet Technology - Source: Archive - Source URL: None TCP is the well-known protocol for reliable transfer of data. But how does it actually work? **System of handshakes ensures reliable transmission of data.** - What happens if a packet is corrupted, and its bits are flipped? - Use a checksum for each packet. - What happens if a packet is lost somewhere? - Wait for acknowledgement from receiver, and if not received, resend. - What happens if a packet is duplicated? etc. etc. There are lots of problems that can happen with UDP. Let's create a protocol that doesn't have this unreliability. Our new protocol has a sequence number for each packet, either 0 or 1. This is known as an *alternating bit protocol*. When it sends a packet, it waits for ACK (acknowledgment) before it sends the next packet. This will take one RTT (round trip transmission) per packet. This is obviously pretty inefficient, but it does solve at least some of these issues. It solves the duplicate issue because if it sends the same packet, it will have the same bit sequence number and hterefore will not be taken into account. You can think of this as an infinite loop which flips its break condition after being broken. This is possible because you're only sending one packet at a time. Here is time and utilization. $$ T_{transmit} = \frac{L_{packetLength}}{R_{transmissionRate}} $$ $$ U_{sender} = \frac{\frac{L}{R}}{RTT - \frac{L}{R}} $$ If we can send W packets at aa time, then we can replace $\frac{L}{R}$ with \\(W\\), which will improve our time by a factor of \\(W\\)! However, if we send a stream of packets, then there are more issues. If you send a receiver too much data, then it will throw out the data it cannot receive. If packets get lost somewhere in the router, you will have no idea. Stopping the stream to check for issues wastes a lot of time. We can solve this using a window time. We send packets up until the first ACK is received, then we check for any problems with ACK. But what is the window size? $$ B \cdot RTT = W \cdot packetSize $$ This way we can pipeline as many packets as are possible before anything can go wrong. This increases the complexity on the recv side, but keeps the connection simple and reliable. Let's say the window size is 3. We can send 3 packets, then we must wait for ACK for the first packet. When the ACK is received for packet p, we send packet p + 3. If the timer expires, we have to resend the packet. We can make the assumption that everything to the left of the window has been ACKed, and everything to the right of the window has not yet been sent. What is the efficiency of this new method? $$ U_{sender} = \frac{3 \cdot \frac{L}{R}}{RTT \cdot \frac{L}{R}} $$ This is an increase of a factor of 3! But what if a packet is dropped and sent later? This is where the receiver comes in. There are two ways of dealing with this: 1. Go back N: simple, keep a buffer of 1 and throw out data with greater sequence numbers and ACK at that packet that was received. This will force the sender to resend all of the packets thata re now timed out. However, this also throws out correctly received packets. 2. Selective Repeat: complex, keep a buffer proportional to W and keep data that has been received, cancelling the timers for each. Once the data has been correctly received, you can move the window and keep the buffer open for anything that hasn't been received. Let's say packet 1 is timed out but 2 and 3 came in correctly. We can cancel those two timers and slide the window to 4 and 5 but still wait for 1. *There are only ever W packets in transmission at any given point!* In order to demonstrate the benefits of selective repeat over go back N, we'll work though an example. Say we are sending 5 packets of size 100 B each over a 100 MB/s link with a RTT of 100 ms, with packet timeout being send_time + RTT and the window buffer being of size 5. How long would it take to receive all of the packets if we drop the third packet? The send_time here is 100 B / 100 MB/s or 0.001 ms. The time to receive a packet is half of RTT or 50 ms. The timeout length is 100.001 ms. With go back N, packets 1 and 2 are received at times 50.001 ms and 50.002 ms. All good! But when the third packet is dropped, we must wait 100.01 ms to receive it again at 150.003 ms. Even though we received packets 4 and 5 already at 50.004 ms and 50.005 ms, we had to throw them out because we didn't get packet 3. We resend them as soon as possible after packet 3 so they are received at 150.004 ms and 150.005 ms respectively. The whole transaction takes 150.005 ms. With selective repeat, the four non-dropped packets are received in the same way. However, after packet 3 times out and is resent, it is the only packet that is sent at 150.003 ms, since every other packet was within our window. The whole transaction here will take 150.003 ms. The protocol we have just created is known as **TCP**. - connection management - retransmission - flow control - congestion control - frame format It has it all! TCP is the default protocol for every web browser in existence and is the de facto protocol for reliable transmission. Chrome now also uses QUIC which is a Google-created protocol but the vast majority of Internet traffic is still TCP. ## TCP Header The TCP header is much more complex than the UDP header. It includes some of the same elements, such as the port numbers and the checksum, and others: - 32-bit sequence number and acknowledgement number, which are used for reliability - 16-bit receive window for *flow control* so that the bytes sent do not overwhelm the receiver - 4-bit header length field to indicate the length of the *header*, needed because of variable options, but usually 20 - variable options field needed when specific maximum segment size (MSS) is needed or other options for high-speed networks - 6-bit flag field: ACK for a segment that is an acknowledgement; RST, SYN, and FIN for connections, PSH for immediate network sending, and URG to indicate some urgent data is located at the location where the 16-bit urgent data pointer points (these last three are typically unused) The sequence and acknowledgment numbers are how TCP messages are ordered despite being sent in a full-duplex way. The sequence number orders the segments for reconstruction into a full message, and the acknowledgment number is a way of communicating the next byte that the sender needs. This way, if a segment is dropped, the sender of the dropped segment knows that the receiver didn't get it because they're still waiting on it in the acknowledgement number. ## RTT Round-trip time is an important concept when we consider making our protocol as fast as possible. This is how long it takes for a segment to be sent and an acknowledgment for that segment to be recieved. In order to get an accurate measure of the RTT, TCP takes the current RTT and refines it to an average as more segments are transmitted, since any one SampleRTT (SRTT) might be an outlier. In order to get the average EstimatedRTT we follow this formula: $$ERTT = (1 - α) \cdot ERTT + α \cdot SRTT$$ \\(α\\) in this formula is typically 0.125. This means that average is weighted towards recent samples, since they better reflect the current weight. We should also have a measure of variance in RTT, since averages can be deceiving. DevRTT is the measure of the variance, which is calculated like this: $$DevRTT = (1 - β) \cdot DevRTT + β \cdot | SampleRTT - EstimatedRTT |$$ Like ERTT, DRTT moves and is weighted recently, with a \\(β\\) value of usually 0.25. The reason this is important is because the interval at which a TCP segment transmission times out is based on RTT. TCP calculates this way: $$TimeoutInterval = ERTT + 4 \cdot DRTT$$ This starts out as one second, then changes over time. Timeouts cause this value to double because a new link may be chosen that has a longer RTT than what ERTT has estimated so far. ## Congestion Control A network is said to be **congested** when so much traffic is being sent through the network that many packets become lost. To combat this, TCP implements **congestion control**. Every sender has a limited rate at which it can send traffic into its connection based on the way TCP thinks the network is congested. If it thinks it's not very congested, then it increases the send rate, and vice versa. There are three elements to this implementation: how to limit the rate, how to perceive the congestion, and how mathematically the rate should be limited. The mechanism at the sender keeps track of a *congestion window*, also known as `cwnd`. The amount of unacknowledged data that a sender sends must not be greater than either the congestion window or the receive window. ## Electromagnetic Induction - URL: https://sharifhsn.dev/blog/induction/ - Structured data: https://sharifhsn.dev/api/posts/induction.json - Description: The final piece to the puzzle of electromagnetism here is electromagnetic induction, where we use magnets to create current. - Date: 2022-02-18 - Exact published timestamp: 2022-02-18 - Topics: Physics - Categories: Physics - Source: Archive - Source URL: None The final piece to the puzzle of electromagnetism here is **electromagnetic induction**, where we use magnets to create current. ## Induced Emf/Current Like with charges in a magnetic field, magnets do not generate current unless they are moving relative to a wire. The changing magnetic field \\(\vec{B}\\) is what creates the current. This current is called an **induced current**, and the "emf" that causes it which is the wire itself is an **induced emf**. The emf can also be induced by changing the area of coil in a magnetic field, as that changes the relationship between the magnetic field and coil. These examples are in a closed circuit, which has a current. An open circuit would not have a current, but it would still have the induced emf. ## Motional Emf An induced emf is created in a metal rod that moves through a magnetic field as long as the velocity is not parallel to the magnetic field. We can use RHR1 to determine where the positive and negative sides of the rod are when the rod has a velocity through a magnetic field. The movement of these charges to either end of the rod cause a charge separation to occur, which leads to a current running from the negative end to the positive end. This is called **motional emf** because it comes from the motion through a magnetic field. We can find the motional emf when the length, velocity, and magnetic field are all perpendicular i.e. they all have their own axis: $$Ε = vBL$$ There is another magnetic force that opposes the motional emf. The current creates its own magnetic field, which by RHR1 would actually oppose the velocity of a mutually perpendicular system. We have to have an external force moving the rod otherwise the magnetic field produced by its own current will cause it to stop. ## Magnetic Flux Remember electric flux? Magnetic flux is quite similar in that it is defined as amount of magnetic field passing through an area, so \\(Φ = BA\cos{θ}\\) in \\(Wb\\) or \\(T \cdot m^2\\). We can define motional emf through magnetic flux like so: $$Ε = -\frac{ΔΦ}{Δt}$$ or, emf is the rate of change of magnetic flux. This is why we can induce an emf by changing the area of a coil. The reason that the equation is negative is because the induced current will create a magnetic force which will oppose its velocity direction. ## Faraday's Law **Faraday's Law** is precisely the equation we just laid out but with one additional component to account for loops: $$Ε = -N\frac{ΔΦ}{Δt}$$ The emf is generated if the flux changes, which depends on \\(B\\), \\(A\\), or \\(ɸ\\), which are magnetic field, area, and angle of the magnetic field with respect to the normal of the surface. ## Lenz's Law **Lenz's Law** is less of an equation and more of a method to understand the polarity of an induced emf. Let's follow the reasoning: 1. Check whether the flux is increasing or decreasing. 2. Find the direction of the induced magnetic field to *oppose* the change in flux. 3. Use RHR2 with the previous step to find the positive and negative end, with the polarity coming out of your palm. This is best understood through an example. Let's imagine we have a loop with a bar magnet approaching at north side head on. 1. The flux is increasing, because the magnet is getting closer. 2. The induced magnetic field must oppose the magnet field to decrease the flux. 3. RHR2 tells us that if the magnetic field is coming out of our palm at the magnet, the current must be going counter-clockwise. Therefore, the left point is positive and the right is negative, as the external circuit must have positive going to negative. ## Transformers A **transformer** is a device that increases or decreases ac voltage. It is responsible for reducing the voltage in your phone charge so you're not sending 120 V to your phone that doesn't need it. This change is called a **step-up** or **step-down** which increases or decreases the voltage, respectively, directly. When we send power over power lines, there is a certain power loss that happens depending on the voltage and the resistance of the transmitting power wire: $$P_{lost} = I^2R$$ which we subtract from our initial power to get the final power. If we want to figure out how much voltage we need to step-up or step-down to get a specific power loss, we just resubstitute the desired power loss into that equation to get the current and proportion it to the voltage. ## Data Types in OCaml - URL: https://sharifhsn.dev/blog/data-types/ - Structured data: https://sharifhsn.dev/api/posts/data-types.json - Description: When we make our programs more and more complex, we need more complex data types as well. We have only used OCaml's built-in data types, how can we construct our own data types? - Date: 2022-02-17 - Exact published timestamp: 2022-02-17 - Topics: Principles of Programming Languages - Categories: Principles of Programming Languages - Source: Archive - Source URL: None When we make our programs more and more complex, we need more complex data types as well. We have only used OCaml's built-in data types, how can we construct our own data types? ## `type` The `type` keyword is similar to the `typedef` keyword in C, except more limited in scope. A `type` can only be multiple variants of arbitrary values. ```ocaml (* coin is enum with variants Heads and Tails*) type coin = Heads | Tails ``` Each variant can also contain data of other data types. ```ocaml type shape = | Rect of float * float | Circle of float let r = Rect (3.0, 4.0) (* r has type shape *) ``` `shape` here has two variants. It can either be a tuple of two `float`s when it is a `Rect`, or it can be a single `float` when it is a `Circle`. These data types are also known as *algebraic data types* or *tagged unions*. ## Option ADTs can be useful when we want to ensure the complete handling of all cases. For example, if an object is nullable, it is useful to make sure that we must handle the null case instead of passing that off to the developer who might carelessly not handle it. This is where the **option** type comes from. ```ocaml type 'a option = | Some of 'a | None ``` The `'a` keyword means that that the type `option` is polymorphic, and the variant `Some` will contain whatever type that `option` is defined for. When handling an `option`, you *must* destructure it into its `Some` and `None` variants and handle both cases, otherwise OCaml will warn you for a non-exhaustive pattern match. ## List We can actually define our own list data type as a **recursive data type**, which is a data type which contains itself. ```ocaml type 'a list = | Nil | Cons of 'a * 'a list ``` Here, `list` has two variants, `Nil` and a `Cons` tuple of an element and a `list`. If we think of the traditional list data type, this is actually just a more verbose version. `[]` is sugar for `Nil` and `::` is sugar for `Cons` tuple. ```ocaml let rec len l = match l with | Nil -> 0 | Cons (_, t) -> 1 + (len t) (* same as *) let rec len l = match l with | [] -> 0 | _ :: t -> 1 + (len t) ``` ## Exceptions **Exceptions** are a special data type used for errors in OCaml. Exceptions are similar to type constructors in that they can take arguments or have none. ```ocaml exception Sign of int let f n = if n > 0 then raise (Sign n) else raise (Failure "foo") ``` We can `raise` an exception with arguments whenever we want, which will exit the function with the exception name and its arguments. `Failure` is a generic exception type that is used with strings. There is also special `try` syntax used to catch exceptions. ```ocaml let g n = try f n with Sign n -> Printf.printf "Caught %d\n" n | Failure s -> Printf.printf "Caught %s\n" s ``` The function `g` will try running `f n`, but if that raises an exception, it will be caught in `with`. It can be pattern matched for different exception types. ## Operational Semantics - URL: https://sharifhsn.dev/blog/operational-semantics/ - Structured data: https://sharifhsn.dev/api/posts/operational-semantics.json - Description: What are the formal semantics of how a programming language works, mathematically? That's a broad topic. Operational semantics, which are how programs execute, are narrow enough fo… - Date: 2022-02-17 - Exact published timestamp: 2022-02-17 - Topics: Principles of Programming Languages - Categories: Principles of Programming Languages - Source: Archive - Source URL: None What are the formal semantics of how a programming language works, mathematically? That's a broad topic. **Operational semantics**, which are how programs *execute*, are narrow enough for one article. ## Rules The basis behind the mathematics of operation are based on using **rules** to define a **judgment**. The expression `e` will always evaluate to the value `v`. We can construct a micro-OCaml through the datatypes `exp` and `value`: ```ocaml eval: exp -> value ``` This way of presenting semantics is called a **definitional interpreter**. `eval` means `exp -> value`, and we use interpretations to define the language's meaning. ## Grammar It is useful to define **meta-variables** which represent categories of syntax. We will list a few here: - `x`: any identifier or variable name - `n`: a numeral value - `e`: any expression - `::=`: meta-syntax for definition - `|`: meta-syntax for variants ```bnf e ::= x | n | e + e | let x = e in e ``` This is an important line which defines what an expression *exactly* is. Up until now, we have used the term expression quite loosely, so it's good to have an exact definition in operational semantics. An expression is either an identifier, a numeral or a combination of expressions. The last variant may seem a little strange to you, but remember, `let` expressions are simply replacing identifiers in an expression, so they are also expressions. This is a powerful definition because it fully encompasses the **abstract syntax tree (AST)**. We can define an evaluation as `eval: exp -> value` where evaluation is the process of turning an expression into a final value (in this case an `int` only). ## Rules of Inference In order to prove the veracity of a judgment, we check its composite rules. For example, let's prove that `1 + 3 ⇒ 4` is true. `1` and `3` are expressions in the sum function which evaluate to their own numeral value. Two values being added together is their sum, which is `4`. The notation of **rules of inference** are used to present rules in formal mathematics: $$\frac{H_1 \mathellipsis H_n}{C}$$ If all the hypotheses are true, then the conclusion is true. If there are no hypotheses, then the conclusion is automatically true (**axiom**). Let's express those same rules about numeral self evaluation and sum expressions in rules of inference: $$\frac{e_1 ⇒ n_1 \quad e_2 => n_2 \quad n_3 \ \text{is} \ n_1 + n_2}{e_1 + e_2 ⇒ n_3}$$ We can similarly describe the more complicated rules of let expressions: $$\frac{e_1 ⇒ v1 \quad e_2\\{v_1/x\\} ⇒ v_2}{\text{let} \ x = e_1 \ \text{in} \ e_2 ⇒ v_2}$$ ## Derivations The **derivation** is a process in which we apply rules to an expression in succession. We take our conclusion, then break it up into its constituent rules. If any of the rules need more rules themselves, we break them up too. Think of it like a tree which expands out from the conclusion. Let's use `let x = 4 in x + 3 ⇒ 7` as an example. $$\frac{4 ⇒ 4 \quad \dfrac{4 ⇒ 4 \quad 3 ⇒ 3 \quad 7 \ \text{is}\ 4 + 3}{4 + 3 ⇒ 7}}{\text{let}\ x = 4 \ \text{in} \ x + 3 ⇒ 7}$$ Let's look at this derivation step by step. First, we start at our conclusion. We say that `x = 4` in the let expression. Because `4` is used an expression here, we need to generate the hypothesis that the expression `4` will evaluate to itself, which is true from our reflective axiom. Then, we need to show that the resulting expression `4 + 3 ⇒ 7` is valid, which requires the sum rule of inference. Knowing that `4` and `3` evaluate to themselves and that `7` is the sum of those two values, then we can confirm that hypothesis and therefore the conclusion! The way that we've written this derivation is recursive in nature. This is how definitional interpreters will evaluate expressions; they will search the expression for any constituent and evaluate them in turn in order to evaluate the entire expression. $$\frac{\text{eval Num } 4 ⇒ 4 \quad \text{Plus(Ident("x"), Num 3)}}{\text{eval Let("x", Num 4, Plus(Ident("x"), Num 3))}}$$ All evaluation is mathematical proof. An expression `e` that provably evaluates to value `v` **is** `v`. ## Environment From a mathematical perspective, an environment is a partial function that maps identifiers to values. Because it's partial, not all possible identifiers are mapped, so some identifiers are undefined. The notation for an empty environment is `•`, which is undefined for every identifier. We can use notation here like in OCaml lists where we can define arguments as either the nil form `•` or a mapping cons the rest of the environment. Lookup of a value is similarly recursive like an OCaml list. In fact, in an OCaml definitional interpreter, the environment has type `(id * value) list` where we can search and match mappings like a regular *association list*. We can add environment semantics to our previous understandings of judgments like so: `A; e ⇒ v`. Now, when we use identifiers in expressions, we can search the environment for the equivalent value and substitute it. The environment is the formal way of representing *state* in an operation. ## Conditionals So far all we've done in terms of operational semantics is define variables and perform simple sums on them. One of the greatest powers of programming is the ability to implement control flow: if this, then that. Since we're programming functionally and not imperatively, these conditionals evaluate to expressions, not just operations, so we must have else as well to account for all cases. To accommodate this, we must expand our earlier definition of expressions to include *equality* and the *if expression*. The if expression takes an equality expression as a conditional, and if it satisfies the boolean argument then the if body is returned as the evaluation, else the else body. # ## Tail Recursion in OCaml - URL: https://sharifhsn.dev/blog/tail-recursion/ - Structured data: https://sharifhsn.dev/api/posts/tail-recursion.json - Description: Recursion in many languages can cause significant overhead. It might seem that the excess amount of recursion in OCaml would decrease its performance. But it actually doesn't, and … - Date: 2022-02-15 - Exact published timestamp: 2022-02-15 - Topics: Principles of Programming Languages - Categories: Principles of Programming Languages - Source: Archive - Source URL: None Recursion in many languages can cause significant overhead. It might seem that the excess amount of recursion in OCaml would decrease its performance. But it actually doesn't, and that's thanks to **tail recursion**. ## Implementing Reverse In a functional language with a linked list, it's not a trivial task to reverse a list. There are multiple ways to do it, and they have different performance implications. The most basic implementation would just be to concatenate backwards. ```ocaml let rec rev l = match l with | [] -> [] | x :: xs -> (rev xs) @ [x] ``` Unfortunately, there's a problem with this. Every time we call `rev` recursively, we must initialize a new stack frame for every recursive call, leaving the value `x` in `[x]` in the calling stack. ```ocaml rev [1; 2; 3] → (rev [2; 3]) @ [1] → ((rev [3]) @ [2]) @ [1] → (((rev []) @ [3]) @ [2]) @ [1] → (([] @ [3]) @ [2]) @ [1] → ([3] @ [2]) @ [1] → [3; 2] @ [1] → [3; 2; 1] ``` As you can see, there are a total of three stack frames for a list with three elements. This is pretty bad. However, we can rewrite this function to use tail recursion. ```ocaml let rec rev_helper l acc = match l with | [] -> acc | x :: xs -> rev_helper xs (x :: acc) let rev l = rev_helper l [] ``` There doesn't need to be any stack frame here because there are no local variables. The only thing that's being returned is the function call to `rev_helper`, so the stack can simply change to the `rev_helper` call without saving the previous stack frame. This tail recursion is incredibly powerful. I'll repeat it to be clear. Tail recursion works when the return value of a recursive function is *only* the recursive function call. The reason that the previous `rev` function didn't work is because the return value was `(rev xs) @ [x]`, so `x` must be saved in a stack frame in order to remember it. However, you might have noticed that we had to include a new variable called `acc`. This accumulator variable is a common pattern for when we want to have tail recursion since the `acc` is located inside the function call as an argument. The power of tail recursion can not be understated here. In a typical function, having excessive stack frames can easily cause a stack overflow for a large data set. With tail recursion, we can have both the lack of side effects associated with recursion and the performance associated with iteration. ## General Tail Recursion Pattern ```ocaml let f x = let rec aux arg acc = if (* base case *) then acc else let arg' = (* next argument *) let acc' = (* updated accumulator *) aux arg' acc' in aux x (* initial value of accumulator e.g. 0, []*) ``` ## UDP Explained - URL: https://sharifhsn.dev/blog/udp/ - Structured data: https://sharifhsn.dev/api/posts/udp.json - Description: The simplest transport protocol that is used communicate between sockets is known as UDP, the User Datagram Protocol. - Date: 2022-02-14 - Exact published timestamp: 2022-02-14 - Topics: Internet Technology - Categories: Internet Technology - Source: Archive - Source URL: None The simplest transport protocol that is used communicate between sockets is known as **UDP**, the **User Datagram Protocol**. ## UDP Let's say we want to design the simplest transport protocol possible. All this protocol needs to do is get a message from an application at the socket and send it over the network to the other socket. However, there is one more function that it needs to perform: multiplexing/demultiplexing. The *segment* (transport layer version of a datagram) must be directed to a specific socket. There's multiple sockets that applications communicate with and they all need to be multiplexed together to create a segment, which must then be demultiplexed for the end host to understand what socket it is sent to. You can think of receiving multiple envelopes in your mailbox, then demultiplexing by reading who each envelope is addressed to and giving it to the member of your household to whom it is addressed. Then, when your household members want to send all of their mail out, you multiplex it into the mailbox. UDP attaches a header with source and destination port numbers as fields for multiplexing, a length field, and a checksum to protect against packet corruption. Each of these fields is two bytes in length, resulting in a tiny 8-byte header. ## Magnets - URL: https://sharifhsn.dev/blog/magnets/ - Structured data: https://sharifhsn.dev/api/posts/magnets.json - Description: The question of how magnets work has [long puzzled many](https://www.youtube.com/shorts/8bhYMnHb5JY). We will endeavor to answer all of those questions today. - Date: 2022-02-11 - Exact published timestamp: 2022-02-11 - Topics: Physics - Categories: Physics - Source: Archive - Source URL: None The question of how magnets work has [long puzzled many](https://www.youtube.com/shorts/8bhYMnHb5JY). We will endeavor to answer all of those questions today. ## Magnetic Fields Magnets work similarly to electric charges where like poles repel, and unlike poles attract. However, magnets do not ever exist in isolation. Every magnet has a north and south pole, while you can have an isolated positive or negative charge. Every magnet has a magnetic field around it, similar to the electric field around electric charges. The direction of the magnetic field is shown by the north pole. Every magnetic field points from south to north; this is why compasses point north. They are based off of the magnetic field of the polarity of the Earth's own magnetic field. You can apply many of the concepts of electric fields to magnetic fields, such as field lines and the way they curve and indicate field direction. ## Magnetic Force Charges experience electric force in a field, and they experience **magnetic force** in a magnetic field. However, a few conditions must be met: 1. The charge must be moving. 2. The velocity of the charge must have a component perpendicular to the direction of the magnetic field. To visualize the second rule, imagine the magnetic force like a buffeting wind around the charge. The wind only has any force if it has something to push on. If you hold a piece of paper parallel to the direction to the wind, it'll barely move because there's nothing to push. > This is such an important point it's in its own section. **Right Hand Rule No. 1 (RHR1)** will tell us the direction of a magnetic force based on the magnetic field and velocity of charge. Your right thumb points in the direction of the charge, and the fingers point in the direction of the magnetic field. The palm faces in the direction of the magnetic force. > > This is only applies to *positive charges*. Negative charges will have the direction of the force as opposite. The magnitude \\(B\\) of a magnetic field is defined in teslas \\(T\\) by $$B = \frac{F}{|q_0|(v \sin{θ})}$$ where \\(F\\) is the magnitude of the magnetic force on test charge \\(q_0\\) with velocity \\(v\\\). ## Motion We have been comparing electric fields and magnetic fields a lot, but in regards to motion through them they are quite different. Electric fields attract towards its directions, while magnetic fields direct in the other perpendicular direction. If left to its devices, a moving charge would perpetually move in a circle remaining perpendicular to the magnetic field. This means that the work done by a magnetic field is also different. In fact, the magnetic force *cannot* do work, it can only change the direction of the particle, not its speed nor its energy. The force required to keep a charge moving in a circle is: $$F_c = \frac{mv^2}{r} or\ r = \frac{mv}{|q|B}$$ ## Currents in a Magnetic Field We've discussed charges moving through a magnetic field; what about a current? A current is just a collection of moving charges, after all. If we have a wire with a current running through it placed between two magnets, then we can consider the direction of the current as the charge direction in RHR1 to calculate the direction of the magnetic force. We can use a bit of a trick to get the calculation of magnetic force for a current. We can rearrange the earlier equation for a magnetic field to get the equation for force. Current is the same thing as charge over time, so \\(\frac{Δq}{Δt}\\). Length of the current is the same thing as velocity multiplied by time \\((\frac{m}{s} \cdot s)\\). If we multiply these, the \\(Δt\\) will cancel out and we will be left with the same expression as \\\(|q_0|v\\)! Our new equation for force is: $$F = ILB \sin{θ}$$ The angle \\(θ\\) here maximizes current when perpendicular and is zero when parallel, just as with a single charge. ## Torque As a refresher, **torque** describes the rate of change of the angular momentum of an object. Current-carrying wires have magnetic force applied to them which causes them to move. If the wire are in a loop, then the magnetic force is applied as torque which causes the loop to rotate. Its resting state is when the normal of the loop is aligned with the magnetic field. You can think of it like a compass needle, which will turn until it reaches its resting state of pointing towards the north pole. To calculate the torque, we get the force of each side of the loop turning, which is half the width and so half the force. Summed together we get: $$τ = NIAB \sin{θ}$$ \\(N\\) here is the number of loops in the wire and \\(A\\) is the area that the loops make. When the loop is parallel with the magnetic field it experiences the greatest torque, and when it is perpendicular it experiences none. The collective expression \\(NIA\\) is known as the **magnetic moment** of the coil in \\(A \cdot m^2\\). Motors operate using this principle of a coil turning in a magnetic field. ## Magnetic Fields in a Current Current-carrying wires create their own kinds of magnetic fields. > The second Right-Hand Rule (RHR2) is easier to remember, as it's just a thumbs-up. The thumb points towards the direction of the current, and the fingers curl in the direction of the magnetic field generated. The magnetic field magnitude is given by the following equation: $$B = \frac{μ_0I}{2πr}$$ with \\(μ_0\\) representing the *permeability of free space* with the value \\(μ_0 = 4π × 10^{-7} T \cdot m/A\\). This equation is for an infinitely long, straight wire, which is not necessarily the case. Because of this property, currents can affect each other magnetically. Currents in the same direction are attrated to each other. Currents in a loop have a slightly different magnetic field equation to the straight wire: $$B = N\frac{μ_0I}{2R}$$ in the center of the loop, where the field is strongest. A useful visual for the magnetic field of a current loop is a bar magnet placed in the middle of the loop. If we use RHR2 here, the north pole is on the palm and the south pole is on the back of the hand. ## Ampère's Law **For any current geometry that produces a magnetic field that does not change in time,** $$ΣB_{||}Δl = μ_0I$$ **where \\(Δl\\) is a small segment of length along a closed path of arbitrary shape around the current, \\(B_{||}\\) is the component of the magnetic field parallel to \\(Δl\\), \\(I\\) i sthe net current passing through the surface bounded by the path, and the \\(μ_0\\) is the permeability of free space. The symbol \\(Σ\\) indicates the sum of all \\(B_{||}Δl\\) terms must be taken around the closed path.** ## SMTP Explained - URL: https://sharifhsn.dev/blog/smtp/ - Structured data: https://sharifhsn.dev/api/posts/smtp.json - Description: Email is the most popular way in the world to send large messages. This is a complicated process which requires its own protocol: SMTP. - Date: 2022-02-10 - Exact published timestamp: 2022-02-10 - Topics: Internet Technology - Categories: Internet Technology - Source: Archive - Source URL: None Email is the most popular way in the world to send large messages. This is a complicated process which requires its own protocol: **SMTP**. ## SMTP The **Simple Mail Transfer Protocol** is how user agents communicate with mail servers. On a high level, email works similarly to physical mail. The mail server is like a post office with mailboxes for each user agent. When a person wants to send a message, they put the envelope in their own mailbox. The mailman user agent then takes the envelope to the post office and puts it in their server mailbox. When a person wants to retrieve a message, their mailman gets all the mail in their post office mailbox and delivers it to the person's personal mailbox. The SMTP post office workers at the server are responsible for moving the mail between mailboxes. To continue the post office analogy, we can imagine an email being sent between Alice and Bob. Alice lives in Hong Kong and Bob lives in San Juan, so they have different post offices. The workers at the Hong Kong post office see that there is an envelope in Alice's message queue addressed to Hong Kong. They open a TCP connection as the client to the San Juan server and sends the envelope to the San Juan post office. The workers at San Juan now have an envelope, and they see that it is addressed to Bob. They then place the envelope in Bob's mailbox. When Bob wants to check his email, he sends his mailman over to the post office, who returns with the email from Alice. Let's examine the TCP connection a little more closely. The connection is established on port 25, which is reserved for SMTP. It does some handshaking on the application layer in order to establish the email addresses of the sender and the recipient. The simple exchange goes like this: - server sends 220 with hostname, client responds with HELO and its own hostname - server confirms 250 that it's ok - client tells server MAIL FROM and RCPT TO addresses and the server confirms each with 250 - client sends `DATA` and server confirms with 354 that it's ready to receive mail - the client sends the entirety of the message, ending with a lone period to finish the message, which the server confirms with 250 - the client can repeat steps 3-5 for any additional messages, then sends `QUIT` for a server 221 response that closes the connection ## vs. HTTP Both HTTP and SMTP are protocols that transfer files between clients and servers, so there are natural parallels. However, there are also significant differences. There is a difference in who initiates the connection which defines the two. In HTTP, the *requesting* client initiates the connection, making it a **pull protocol**; you can think of the client "pulling" the data from the server. In SMTP, it is the *sending* client that initiates the connection, making it a **push protocol** where the client "pushes" the data to the other server. SMTP also has many restrictions. The message can only be in ASCII, so any data that is sent over that is not in ASCII, like multimedia, must be encoded and decoded. HTTP allows any data to be transferred. A result of this is that HTTP data of different types like multimedia must be in its own response message, while SMTP objects are all in the same ASCII message. There is some peripheral data that SMTP messages can include in a header, just like HTTP, such as the sender, receiver, and the subject of the email. *These are different from commands; this data is part of the actual email, not the protocol!* ## Mail Access There's a part of the post office analogy that has been ignored so far. How exactly does the mailman get to the post office? After all, this is another client-server connection, so there must also be a protocol that defines this interaction. We can use SMTP for the mailman delivering the envelope, as this is the same kind of "push" that the post office uses to send to the other post office. But this introduces an issue for the person on the other end; if the mailman can only push mail it already has to somewhere else, then how does Bob's mailman get the mail on the server? There are many protocols that can define this interaction, such as **POP3**, **IMAP**, and HTTP. In 2022, HTTP is the obvious choice for this exchange. After all, everyone accesses email through their web browser, which is making HTTP requests anyway. Why not just use an HTTP request to get the mail from the mail server? And as you might expect, almost every email provider today, from Gmail to Hotmail to Yahoo Mail uses Web-based email. However, back when "user agent" and "web browser" weren't synonymous, people had applications dedicated to accessing email that couldn't access the web or use HTTP. In those times, **Post Office Protocol—Version 3** and **Internet Mail Access Protocol** were used. POP3 is a simple and therefore limited protocol. It operates over port 110 and has three steps: authorization, transaction, and update. The post office first authenticates the mailman by asking him for a username and password before he can check your mailbox. Then, the mailman retrieves the envelope from the post office. While the mailman is at the post office, they can also do some operations like marking emails for deletion and getting statistics for your account. When the mailman leaves the post office, the workers throw away all the mail that the mailman marked for deletion. It's common for users to want to organize emails through folders like Inbox, Junk, etc. This is of course possible through user agent, but POP3 does not allow for this organization in your post office mailbox. IMAP provides this as well as many more features. It is stateful so that the user can keep the folders that it creates at the mailbox. One more important feature is its ability to lazily fetch emails from the mailbox, so the user's entire inbox is not just sent immediately. Imagine if you had to download every email in your inbox in order to access it! I know my connection wouldn't be able to handle my 1k+ unread emails. ## Higher Order Functions in OCaml - URL: https://sharifhsn.dev/blog/higher-order-functions/ - Structured data: https://sharifhsn.dev/api/posts/higher-order-functions.json - Description: We can do a lot of cool things with functions besides calling them in OCaml. - Date: 2022-02-08 - Exact published timestamp: 2022-02-08 - Topics: Principles of Programming Languages - Categories: Principles of Programming Languages - Source: Archive - Source URL: None We can do a lot of cool things with functions besides calling them in OCaml. ## Anonymous Functions Values are a subset of expressions, as previously stated. All expressions can evaluate to values, but values are final. **Anonymous functions** are also values. Sometimes, it's more convenient not to create and name a whole new function for our purpose. Anonymous functions are ad hoc functions that exist as values in expressions. They are expressed using the keyword `fun`. ```ocaml let y = fun x -> x + 3 ``` This might not seem to have much benefit compared to a full function definition, but it is very useful within `let` expressions. Since anonymous functions are values, not just expressions, they can be manipulated far more powerfully than even general expressions. ```ocaml let y = (fun x -> x + 1) 2 in (fun z -> z - 2) y ``` This code might seem a little hard to parse, but it's easier to think about if we rewrite it to use traditional function definitions. ```ocaml let f x = x + 1 let g z = z - 2 let y = f 2 in g y ``` Now we can tell that `y` is 3 in the function `g`, which then evaluates to 1. However, the former code snippet is a much terser way to write this expression if we don't need the functions `f` and `g` anymore. One good way to think about it is that anonymous functions are to regular functions as literals are to variables. If we only need to use the value `"really_long_string"` once, we don't need to store it in a variable. On the other hand, it can be useful to store that literal in a variable `s` that is much shorter to write. Similarly, if we only need to use the function `x -> x + 1` once, we don't need to store it in a function variable. In fact, this isn't even an analogy. Functions are first-class in OCaml, so regular functions are just variables that store anonymous functions: ```ocaml let f x = body (* this is sugar for this *) let f = fun x -> body ``` And in the same vein, we can name functions within `let` expressions in an anonymous ways. ```ocaml let move l x = let left x = x - 1 in let right x = x + 1 in if l then left x else right x ;; (* same as *) let move' l x = if l then (fun y -> y - 1) x else (fun y -> y + 1) x ``` Note also that the local variable in the anonymous function doesn't actually matter to the expression it's used in; this is a consequence of the shadowing rules of OCaml. There are several functions in the standard library of OCaml that use higher order functions. ## Map `map` is a function in the `List` module of OCaml. Like the name implies, this function maps a function onto every element of a list and returns that list. It has type `('a -> 'b) -> 'a list -> 'b list`. ```ocaml let rec map f l = match l with | [] -> [] | h :: t -> (f h) :: (map f t) ``` This is a simple, yet powerful and useful function. That's why it is included in the `List` module, although it's trivial to write yourself. ## Fold `fold` is another function in the `List` module in OCaml, that iterates over a list. The essential idea is that you have an accumulator variable that you want to get based on the values in a list. ```ocaml let rec fold f acc l = match l with | [] -> acc | h :: t -> fold f (f acc h) t ``` For example, this is a way to implement a sum function using `fold`. ```ocaml let rec sum acc l = match l with | [] -> acc | h :: t -> sum (acc + h) t sum 0 [2; 5; 100; 53];; (* same as *) fold (fun acc x -> acc + x) 0 [2; 5; 100; 53];; ``` Its type is `('a -> 'b -> 'a) -> 'a -> 'b list -> 'a`. We can deconstruct that and understand each part of the function. The initial function `f` applies the type `'b` to `'a` and returns `'a`. Then for the next two arguments, we keep the same `'a` accumulator and iterate over the `'b list`. The use of an accumulator makes `fold` very versatile, since you can put anything in there. We can combine `map` and `fold` to create the *map/reduce* framework which can be massively parallelized. We first map a function over our list, then we reduce the list into a single accumulator value. There is also an alternative version of `fold` called `fold_right` that works in reverse, which can be better for certain problems. However, it comes with steep performance cost: every recursive call builds a new stack frame. The original `fold` is able to optimize this call away by using tail recursion. ## FTP Explained - URL: https://sharifhsn.dev/blog/ftp/ - Structured data: https://sharifhsn.dev/api/posts/ftp.json - Description: Before we had Google Drive, users needed a way to access files from other people quickly and easily. The method for that was FTP. - Date: 2022-02-07 - Exact published timestamp: 2022-02-07 - Topics: Internet Technology - Categories: Internet Technology - Source: Archive - Source URL: None Before we had Google Drive, users needed a way to access files from other people quickly and easily. The method for that was **FTP**. ## FTP The point of the **File Transfer Protocol**, was, as is obvious, to transfer files between two hosts. Like HTTP, FTP operates between a server and client, each of which have their own filesystem that they want to access. It also creates a TCP connection which has certain verification steps to establish a connection that allows the client to copy files stored in the server or vice versa. Unlike HTTP however, FTP operates using *two* parallel connections, the **control connection** and the **data connection**. As the names imply, the control connection is used for verification and command information such as which file to get in which directory, and the data connection is the connection used to send the files. This is called an *out-of-band* connection, unlike HTTP and SMTP, which have the request information in a header in the same connection and are therefore *in-band*. The session begins with a control connection initiated by the client to identify the user and commands. The server receives this information, and if it approves, it initiates a data connection with the client. Like HTTP 1.0, this is a non-persistent connection that is only open for the passing of a single file. The nature of FTP authentication means that FTP must maintain state about the user account, unlike HTTP which is totally stateless by design. This introduces constraints on the number of simultaneous FTP connections. ## Commands FTP commands work like HTTP commands in request headers. They are ASCII and each lines ends with `\r\n`. Some example commands are: - `USER username` - `PASS password` - `LIST`: works like the `ls` command for the remote directory, sent over data connection - `RETR filename`: retrieves the file from the remote directory, over data - `STOR filename`: stores the file from current directory into remote, over data Every command has a corresponding reply with a status code, similar to HTTP status codes. ## Memory Virtualization - URL: https://sharifhsn.dev/blog/memory-virtualization/ - Structured data: https://sharifhsn.dev/api/posts/memory-virtualization.json - Description: Accessing physical memory can cause big issues if we do it directly. How can we virtualize the memory like the CPU so we can use it efficiently and safely? - Date: 2022-02-07 - Exact published timestamp: 2022-02-07 - Topics: Operating Systems Design - Categories: Operating Systems Design - Source: Archive - Source URL: None Accessing physical memory can cause big issues if we do it directly. How can we virtualize the memory like the CPU so we can use it efficiently and safely? Let's think of a super simple one process system. Our OS is located in the lowest part of memory, and everything after that is reserved for the user. Within an address space, we have the stack-heap structure discussed before. There's a big problem, that processes can access the memory that the OS is stored in and modify it. That's not good! If we want multiple processes as well, we want to give processes memory that is managed by the operating system so they do not interfere with each other. Every process should have a self-contained *address space*. We need dynamic allocation from memory as well so that our programs can become very powerful. The stack is not always sufficient for the programs we want to write. The dynamic heap is where all this memory management magic (say that five times fast) happens. Memory accesses are built into `movl` instructions in the ISA. There are special operations to dereference the location stored in a register and access memory. These are known as *load/store* operations. Memory accesses also happen from the code area by the instruction pointer, though that's typically not shown. If we want to count memory accesses, there is one fetch for every instruction, and another if the instruction is a load/store. ```nasm 0x10: movl 0x8(%rbp), %edi 0x13: addl $0x3, %edi 0x19: movl %edi, 0x8(%rbp) ``` In this example, there are 5 total memory accesses. There is one for every instruction, and each `movl` is a load/store because one of the operands dereferences a memory pointer. Typically, we want to reduce this as much as possible because memory access is really slow, and we can't always rely on the cache. So, what are some strategies to virtualize memory? ## Time Sharing When we virtualized the CPU, we gave each process the illusion that it had its own CPU that were all running at the same time. The way this was done was through context switches that preserved CPU state in memory. We could do something similar by saving memory to disk when the process isn't running. However, there is an immediate problem with this: it would be incredibly slow. Disk I/O even with SSDs is several orders of magnitude slower than DRAM, which is itself orders of magnitude slower than cache/register access. Considering how much processes play around with memory, this would be totally infeasible. ## Static Relocation This is an interesting solution. The instructions stored as static data refer to specific memory locations that are hard coded in the application. However, what we could do is change those pointers and memory to memory that is currently available every time we load the process into memory. For example, imagine the previous example which has instructions at `0x10`, `0x13`, etc. We could imagine that those memory locations are no longer available, so the OS changes the static code portion to `0x1010`, `0x1013`, etc. This means that all `jmp` and load/store instructions would also have to be rewritten as well since they directly refer to memory locations. However, this translation is pretty expensive, although it is simple to implement. There are also security concerns, as always. There is no reason that the new memory couldn't be located somewhere that it can't go. ## Dynamic Relocation We want the power of static relocation, but we need a way to manually protect each process from each other. This is such an important issue that it is usually provided as a hardware component: the **Memory Management Unit** (MMU). Whenever a process generates a virtual address in its address space, the MMU takes on the job for translating that address into a real address and giving it back to the process. This is a good example of modularization; we don't want to have the process to worry about memory, so we offload that job onto a specialized unit designed for that purpose. The MMU is managed only by the OS, so the user can never access it. There are two general operating modes, as discussed previously: kernel space, and user space. Privileged kernel operations have full power over all of memory and the MMU, and users have to ask the OS for memory through virtual translation. ## Address Translation To start out with, we will make some assumptions about how address spaces are laid out. They are all contigous spaces in memory of the same size that are smaller than physical memory. As we will see later with segmentation, this assumption will not hold up, but it useful for now. In order to translate a *virtual address* to a *physical address* in memory, we need some kind of reference space in memory. This is given by the **base** and **bounds** registers, which give the physical memory locations of where the virtual address space starts and ends, respectively. These registers are not part of the regular ISA and are instead part of the previously mentioned MMU, and operating on them requires privilege. The MMU can also throw an exception upon illegal memory access which is handled by our OS. Our OS can manage memory in this simple way through the mechanism of a *free list*, which is a linked list which indicates free spaces, which is also used for `malloc`. As discussed before, the address space has a structure where the stack grows downward and the heap grows upward. However, if you look at a diagram of this for more than five seconds, you might notice a giant chasm between the stack and heap. Relying on our assumption of a contiguous space of memory, this is a lot of wasted memory. Some address spaces fix this problem through **segmentation**. ## Segmentation As the name implies, a segmented address space is split into multiple segments with its own base/bounds pair to indicate its logical beginning and end. We can split every piece of the address space into its own segment, so stack goes in one segment, heap goes in another, etc. This way, although the virtual address space has this nice alignment, we are not wasting physical memory with our *sparse* address space. In order to translate a virtual address, we treat the address as an offset with a segment. For example, if we were trying to access a virtual address in the heap, this would be the formula: $$physicalAddress = physicalBase + (virtualAddress - virtualBase)$$ Before we allow this translation to take place, however, we must check that the value does not exceed the physical bound. If this is the case, then we have caused the infamous **segmentation fault**. A faster way to do with this masks is through bitwise operations. The top two bits of the address refer to the segment, either the code, stack, or heap. The rest of the address is the offset within the segment. This way, we can perform the bounds check before accessing the physical address by checking if the expression in the parentheses exceeds the virtual bounds. However, one issue with this means that each segment gets the same maximum size, which is \\(2^{offset}\\). If we want a bigger heap, we're out of luck. We can solve this by tracking instructions instead of bits. For example, an instruction fetch for `%rip` will come from the code segment, so we don't need to put that in the address that it is in the code segment. Another issue is that the stack segment grows downwards instead of forwards, which means that its segment works differently. We need an extra bit for segments that indicates whether they grow forwards or backwards, which will be set to 0 for the stack and 1 for everything else. If that bit is false, then we subtract the offset from the base instead of adding it as above. ## Sharing :) Sometimes, processes share the same code. It seems wasteful to copy the same code segment every time we have a new process, so modern operating systems implement **sharing** for code segments. Sharing can be implemented for any segment, but it is most common for code since it's read-only. This is important for dynamically linked libraries since many processes will access the same library, like `libc`. Hardware adds extra *protection bits* to each segment indicating its `rwx` value similar to a file. If a segment is indicated to be `r--`, then the OS can secretly share the segment between multiple processes, assured that they will only read it. This concept of a read guard allowing for multiple shared references as opposed to a write guard which only allows for one mutable reference will become *very* important when we discuss concurrency. We have only been working on a few segments so far, but segmentation could theoretically be extended to as many segments as you want. In order to have *fine-grained* segments, a segment table is needed to quickly access thousands of segments. There are a few more things that the OS needs to do in order to support segmentation. One is that it must preserve all base-bounds pairs upon a context switch, since the location of the address space is now more complicated. Another is that the growing and shrinking of segments must be managed through `sbrk`-like system calls that shift the bounds of the heap/stack. The final and most important issues that allocating variable size address spaces runs into the same problems of `malloc` where external fragmentation is difficult to avoid. Like with `malloc`, there is no perfect solution, ranging from algorithmic free lists to compact memory which rearranges segments every time a new one is created. ## Lets, Tuples, and Records in OCaml - URL: https://sharifhsn.dev/blog/lets-tuples-records/ - Structured data: https://sharifhsn.dev/api/posts/lets-tuples-records.json - Description: Lists are the most basic data structure in OCaml, but there are others that we need to be aware of. - Date: 2022-02-03 - Exact published timestamp: 2022-02-03 - Topics: Principles of Programming Languages - Categories: Principles of Programming Languages - Source: Archive - Source URL: None Lists are the most basic data structure in OCaml, but there are others that we need to be aware of. ## Let Expressions We have seen the keyword `let` used to define expressions and store values. However, the same keyword can be used to create expressions which bind variables in other expressions. The `let` *statements* we used before do not evaluate to any value, while **let expressions** do evaluate. ```ocaml let x = 5 in x * 3 ``` The `in` keyword gives us a clue on what's going on. You can think of it almost like a function where we replace variables in the inner expression with the values in the outer expression. The expression above evaluates to 15, because it is `x * 3` *in* which `x` is 5. We can type check this expression where `x` has the same type of the binding expression. If you omit `in`, you can think of that as a `let` expression which is bound in the global scope instead of the scope of the body expression. ```ocaml let x = 37;; ``` In this statement `x` is defined as 37 in the global scope and it can be used elsewhere. I've used the word *scope* a lot here, so let's define that a little more concretely. In the above `let` expression, the variable `x` is not visible in any other part of the program. Let's imagine that we have both of these lines together, the `let` expression and the `let` statement. They're both named `x`, so what would the value be? We can imagine evaluating expressions right to left, and upon encountering a variable, act like it is a pointer to an expression in the outer scope. In the innermost scope, the expression is `x * 3`. This expression has no meaning because `x` is a variable, so we back up one scope and check if `x` is defined. And in fact, it is! So we will replace the `x` in that expression with 5. Note that even though `x` is defined as 37 in the global scope, the inner scope **shadows** the global scope. Shadowing refers to when a variable name is rebound in an inner scope to have a different meaning. Some languages, such as Java, do not allow you to do this because of possible confusion. However, it is sometimes useful to use the same name for different things, so languages like C and OCaml permit shadowing. You can also use `let` expressions inside a function, and this is often good style because it clarifies constants: ```ocaml let area d = let pi = 3.14 in let r = d /. 2.0 in pi *. r *. r ``` Much better than C `#define`, right? OCaml does not permit you to mutate variables. However, you can simulate this by shadowing a variable with a new value: ```ocaml let x = 0;; x = x + 1;; (* not allowed! *) let x = x + 1; (* allowed, but discouraged *) ``` This is kind of an ugly hack so you should avoid it in real code, though it is technically possible under OCaml's rules. We can nest `let` expressions, but this is generally bad practice like shadowing. Realistically, it's usually better to just write linear expressions. `let` expressions don't just have to use a plain variable. We can also use patterns to bind expressions, and if the binding expressions fails to match the pattern then we have an exception. This is useful when we want to extract a particular value from an expression. ```ocaml let [x] = [1] in 1 :: x (* evaluates to [1; 1*) ``` ## Tuples Tuples represent collections, like lists, but they contain a fixed amount of values. The tradeoff is that they can be *heterogenous*, which means they can have multiple types. The type of a tuple is the type of each of its component, separated by asterisks. ```ocaml (1, 2) (* int * int *) (1, "string", 3.5) (* int * string * float *) ``` Because each tuple has a distinct type, a list of tuples can only have one type of tuples in it. Tuples lend themselves particularly well to pattern matching. Instead of having multiple function arguments like is typical in OCaml, we can have one argument that is a tuple and then pattern match it in order to destructure it. This is also a convenient way to return multiple variables from a function, which is otherwise not allowed. ## Records Each element of a tuple is referenced by its position. Sometimes, we want to reference elements by name, like in a dictionary. For this use, we use **records**. Records are a distinct type that must be pre-defined before being used. ```ocaml type date = { month: string; day: int; year: int } ``` Now, we can construct records by using the same brace notation but giving each name a value. ```ocaml let today = { day=3; year=2022; month="f"^"eb" };; ``` You might notice that this has a similar syntax to C-style structs. The fields can be accessed through `.` syntax as in `today.day`, and the order of the struct assignment doesn't matter. Records can also be conveniently destructured, like all other data structures. ## Circuits - URL: https://sharifhsn.dev/blog/circuits/ - Structured data: https://sharifhsn.dev/api/posts/circuits.json - Description: Circuits are the method by which all electricity travels throughout the world and power every single device in existence. It is therefore crucial to understand them while understan… - Date: 2022-02-01 - Exact published timestamp: 2022-02-01 - Topics: Physics - Categories: Physics - Source: Archive - Source URL: None Circuits are the method by which all electricity travels throughout the world and power every single device in existence. It is therefore crucial to understand them while understanding electricity. ## Electromotive Force (emf) and Current The **emf** is similar to voltage, but not quite the same. It is the maximum voltage of a battery, signified by \\(Ε\\) (capital epsilon). It represents the maximum amount of joules of energy that can transfer from a battery based on the charge. **Current** is formed when an electric field parallel to a wire moves free electrons from the positive terminal of a battery to the negative terminal. You can think of it mentally like the flow of charge over time, like water going through a hose, and in fact the equation is \\(I = \frac{Δq}{Δt}\\). For a non-constant flow, this is the average current. Current uses the unit \\(A\\) for ampere. A **direct current (dc)** has electrons that go through the circuit in the same direction at all times, which is negative to positive. An **alternating current (ac)** alternates direction constantly from moment to moment. > *Very important note:* electron flow and conventional current \\(I\\) are ***not the same thing***. Conventional current actually goes opposite to electron flow for historical reasons. When we draw current over circuits, we use conventional current notation, so current flows from positive to negative. This is because it was once thought that circuits worked by the passing of positive charges, not negative charges. ## Ohm's Law **The ratio \\(V/I\\) is a constant, where \\(V\\) is the voltage applied across a piece of material (such as a wire) and \\(I\\) is the current through the material:** $$\frac{V}{I} = R = constant | V = IR$$ **\\(R\\) is the resistance of the piece of material in ohms (Ω).** Again using our hose analogy, resistance is like how narrow the hose opening is which lets our water through. Resistance is typically discussed in the context of **resistors**, which are electrical devices that apply resistance to a circuit. ## Resistance and Resistivity The resistance of a piece of material is dependent on certain qualities. If you think in your head of a small plastic tic-tac as our resistor, this is the formula: $$R = ρ\frac{L}{A}$$ where \\(L\\) is length, \\(A\\) is cross-sectional area, and \\(ρ\\) is a constant called **resistivity** that is constant to a material. Conductors have very low resistivities, and insulators have very high resistivities. Resistivity also typically depends on temperature, though in more complicated ways than can be expressed in the above formula. The most common way to express it as: $$ρ = ρ_0[1 + α(T - T_0)]$$ You can also use resistance instead of resistivity here. The important constant here is \\(α\\), which is the *temperature coefficient of resistivity*, which can be positive or negative depending on how the specific material relates resistivity to temperature. > Some materials, known as *superconductors*, have resistivity of zero at certain temperatures. This means you can circulate a current indefinitely within that circuit without needing a supply of emf from a battery. They can be used in MRI, maglevs, and computer chips. ## [POWER](https://www.youtube.com/watch?v=chPDTUjnWgA) **Power** is an important concept when we want to give energy to electrical appliances. The easiest way to think about it on a high level is the change in energy per unit time. Another way to think about it is the flow of voltage transferred. $$P = \frac{Change\ in\ energy}{Time\ interval} = \frac{(Δq)V}{Δt} = \frac{Δq}{Δt}V = IV$$ The units of power are in watts \\(W\\), which you might remember from lightbulbs. Current multiplied by voltage makes sense; it's the flow multiplied by the quantity, which tells you change in quantity over time. You can also write it as: $$P = IV = I^2R = \frac{V^2}{R}$$ which is often useful when we don't have one of those elements. ## Alternating Current Most batteries in the world use ac instead of dc. Therefore, we must note the differences between it and dc. The voltage is not always the same, and fluctuates constantly according to this equation: $$V = V_0 \sin{2πft}$$ where \\(V_0\\) is the maximum voltage and \\(f\\) is the frequence of isolation. This value is in radians when the sine function is applied. Current oscillates at the same rate, and so does power. We can simply substitute the above equation for \\(V\\) to evaluate those. We often want to understand average current and average voltage in relation to power, which are calculated by dividing the maximum of each by \\(\sqrt{2}\\). We can use these two values to get average power through the analogous power equations. ## Series Wiring All the circuits we've discussed have one device with a straight wire. Often we want to have multiple devices on the same circuit. When we have multiple resistors in series, their resistances added up to get the total resistance of a circuit. A very useful method to solve these kinds of problems is to "combine" resistors by adding up their resistances. This becomes helpful when we start working with parallel wiring. Another way we can think about voltage in this context is that we have voltage that goes through each resistor which is added up to get the final voltage. This will be relevant later when we talk about voltage drops. ## Parallel Wiring Parallel wiring is wiring where the voltage across each device is the same. Now, this might not necessarily make sense; wouldn't the voltage be divided still? The element that is divided in parallel is actually the *current*, whereas in series it is the current that is preserved and the voltage that is divided. This is useful when we want multiple appliances connected to the same power source that do not interfere with each other with respect to voltage. Parallel resistors, paradoxically, act as if they have *less* resistance. Instead of adding up the resistances straightforwardly, parallel resistances add up as so: $$\frac{1}{R_p} = \frac{1}{R_1} + \frac{1}{R_2} + ...$$ A more convenient way to calculate this is $$R_p = \frac{R_1 \cdot R_2}{R_1 + R_2}$$ for only two circuits, though it gets more complex for more. When we have multiple parts in series and parallel, the way to solve comes by combining different series and parallel parts together. Mind that the combinations must be *only* series or *only* parallel. The easiest way is to start furthest from the capacitor and work your way there. ## Internal Resistance We've separated devices into groups of resistors and conducting wire. However, this is a false dichotomy. All materials have some amount of **internal resistance** which resists current. Typically, batteries and wires have extremely low resistance to facilitate current. ## Kirchhoff's Rules If resistors are in series or parallel, we can combine them fairly easily. There's trouble when we can't do that, however. We must use **Kirchhoff's Rules**, specifically the **junction rule** and the **loop rule** to calculate resistance in these cases. The junction rule states that the total current directed into a junction must equal the total current directed out of the junction. This is just an extension of conservation of electric charge that we discussed earlier. You can imagine an intersection of cars where you have four cars waiting at the light. They can either turn right or go straight. Either way they go, however, there must be four cars that are leaving the intersection unless something has gone horribly wrong. The loop rule relates a similar idea but in relation to electric potential. The voltage rises at every capacitor, and it must have drops through every resistor that equal the rise through the resistor. For example, after a 12 V rise through a capacitor, a pass through a 5 Ω and a 1 Ω resistor must drop 10 V and 2 V respectively in order to have it be the same. Some important notes: - the choice of direction is arbitrary, we will simply end up with a negative current if it is wrong - we must indicate the positive and negative directions of every capacitor/resistor in order to find our voltage drops/rises - our final equation will have the drops on one side and the rises on the other side, with every term either being voltage or a resistance multiplied by the variable \\(I\\) - the answer is \\(I\\) which is solved for through normal algebra > In laboratory, there are various devices we use to measure current and voltage. You should remember that current is only the same in series and voltage is only the same in parallel. Therefore, ammeters must be connected in series and voltmeters in parallel. ## Capacitors in Series/Parallel Capacitors work opposite to resistors when they are in series and in parallel. They add up charge simply when placed in parallel, and have the inverse reciprocal relationship when placed in series. The reason this happens is because charge and capacitance are directly related while current and resistance are inversely related. One note is that with regards to the charges on their plates, capacitors in series always have charges of the same magnitude because they will equalize out. ## RC Circuits We often have circuits that have both resistors and capacitors. In this situation, the capacitor's time to charge depends on the resistance: $$q = q_0[1 - e^{-t/(RC)}]$$ \\(RC\\) here is very simply, the resistance multiplied by the capacitance, which ends up being a unit in seconds known as \\(τ\\), the time constant. That is the amount of time for the capacitor to charge to 63.2%, which is \\(1 - e^{-1}\\). Discharging works with a similar equation: $$q = q_0e^{-t/(RC)}$$ where \\(τ\\) represents losing 63.2% charge. ## DNS Explained - URL: https://sharifhsn.dev/blog/dns/ - Structured data: https://sharifhsn.dev/api/posts/dns.json - Description: We need protocols to define the way that messages are interpreted. - Date: 2022-01-31 - Exact published timestamp: 2022-01-31 - Topics: Internet Technology - Categories: Internet Technology - Source: Archive - Source URL: None ## Protocols **We need protocols to define the way that messages are interpreted.** Messages are different based on what kind of information you are sending. For example, *HTTP (HyperText Transport Protocol)* defines the way that webpages are served. Otherwise, the bytes that are sent between a client and a server are meaningless. ## Identification **We can identify hosts using IP addresses and ports.** With cell phones, we use telephone numbers to identify each other and call each other. Sometimes, we can store names in contacts associated with the phone numbers that are easier to remember than unique numbers. *IP addresses* work the same way. Each IP address uniquely identifies a host then can send and receive information through a network. IPv4 is a 32-bit number that was traditionally used, however due to its limited scope (only $2^{32}$ ≈ 4 billion possible IP addresses). IPv6 is the new type of address that is 128-bit number with $2^{128}$ possible addresses. IPv4: `128.6.24.78` IPv6: `2001:4000:A000:C000:6000:B001:412A:8000` There can be more than one application running on a host, however. They can't interfere with each other, so we need a way to distinguish them. This is done via *port numbers*, which are 16-bit numbers that can be bound to applications that communicate over the network. Some port numbers are reserved for special purposes, such as port 25 for email, port 80 for HTTP, port 443 for HTTPS. Port numbers are like different employees that work at a call center. Both client and server must identify each other's IP address and port numbers: a 4-tuple: $(S_{IP}, S_{P\\#}, D_{IP}, D_{P\\#})$ This tuple is a *uniquely defined bidirectional connection*. The OS manages a simple data structure of port numbers to track which are in use, no complexity involved. ## Client-Server Architecture **Many clients communicate with one always-on server.** A *server* is a host that is always on and has a permanent IP address that is often public. For large servers, it is often necessary to use server farms that distribute servers geographically for faster communication across the globe. A *client* is one of many hosts that makes a connection with the server to access data on the server. They might connect intermittently or have dynamic IP addresses. Clients do not communicate with one another directly, instead using the server as a medium. For example, a website server might listen on a fixed, public IP address at port 80, waiting for a HTTP request. A client might request a webpage from this server and make a connection with its own IP address and port 80. Many clients will all have different connection tuples because of their own unique IP addresses. This way, many clients can make connections without confusing the server. One domain name can map to different IP addresses. This is how *google.com* can be split across thousands of servers that serve billions of requests every day to the exact same domain. The domain name can be thought of as a contact with saved name that might have many phone numbers. ## DNS **Domain names are converted to IP addresses through the DNS.** *DNS* is the *Domain Name System* which acts as a service to resolve IP addresses from a normal alphanumeric domain name, just like a telephone book. The DNS is itself an online service that must be accessed through a network. The DNS listens on port 53 and the IP address is well-known. A simple DNS might be a large centralized database of names and IP addresses, around 4 billion for IPv4. However, if this DNS crashes, that would destroy the whole Internet. Also, everybody on the Internet would be trying to access this server, which place tremendous load. It would be a big security issue and a large target for attack, both physical and virtual. Lookup in a server with billions of entries would be slow. Latency for hosts that are physically far away would be greatly increased depending on the server location. Every new host would need to be entered in the same location, which would be slow. *This does not scale!* Real DNS is implemented as a distributed service across different countries. *Top-level domains* like `.com`, `.org`, `.edu`, are most popular and are the main domain names that people access. Each country also has it's own top-level domain server named after it, like `.de`, `.uk`, `.be`. Many websites nowadays utilize country codes for special names that include the domain name e.g. `youtu.be`. Second-level servers are the most commonly accessed domain names, like `amazon.com`, `google.com`, etc. These names are separated by periods, from right to left. Larger domain names like universities can be further namespaced, like `cs.rutgers.edu`. Rutgers University has its own DNS server that manages these additional domain names. The final name that contains the actual data is called the `authoritative` domain. ## Packet Switching vs Message Switching - URL: https://sharifhsn.dev/blog/packet-switching/ - Structured data: https://sharifhsn.dev/api/posts/packet-switching.json - Description: How do we get the best performance in sending information over the internet? What is packet switching and how can it help? Performance is bottlenecked by propagation delay and tran… - Date: 2022-01-31 - Exact published timestamp: 2022-01-31 - Topics: Internet Technology - Categories: Internet Technology - Source: Archive - Source URL: None How do we get the best performance in sending information over the internet? What is packet switching and how can it help? **Performance is bottlenecked by propagation delay and transmission time.** Propagation delay is dictated by the physical distance between the client and server. If you were to send a message to server on Mars, there would be 200 million km to travel across, aka 1000s at the speed of light! New York to Los Angeles is about 20ms at this same speed of 5μs/km. Transmission time is dictated by bandwidth, which tells you how many bytes per second you can send. 1MB/s link speed will take 1ms for 1000B. The reason this happens is because as soon as the first bit is propagated, all the rest of the bits will follow which will be concentrated by bandwidth, 1μs after each other with this link speed. For small messages in the tens or hundreds of bytes like handshake acknowledgements, propagation delay is very significant. For larger messages, propagation delay is ignored and bandwidth is considered more significant. $RTT = 2PD$: Round-trip time is equal to twice the propagation delay. **Packet switching is faster because it is a continuous stream.** Assume 1000B/s link speed between sender and destination. If we do message switching, it is slow. A 1000B would take 1s (duh) to go from sender to router. Similarly, it would take another second to go from router to destination, so 2 seconds in total. Let's assume we split our 1000B message into 100B packets. The first packet will take 0.1s to arrive to the router and another 0.1s to arrive at the destination, so 0.2s. Each packet follows directly after the other, so the second packet takes only 0.1s total because it's already at the router by the time the first packet is at the destination. In total, the message takes 1.1s to travel. These benefits compound when messages are passing over multiple routers, because only the first packet has to deal with the overhead of the link speed for all the routers. All the rest of the packets will only take the time for the link speed between the last router and the destination. However, bottlenecks can still occur when link speeds between routers are different. The weakest link will bottleneck the connection even in packet switching. ## Scheduling - URL: https://sharifhsn.dev/blog/schedulers/ - Structured data: https://sharifhsn.dev/api/posts/schedulers.json - Description: The way that the operating system decides which processes to run and when is a complicated process known as scheduling. This is how it works. - Date: 2022-01-31 - Exact published timestamp: 2022-01-31 - Topics: Operating Systems Design - Categories: Operating Systems Design - Source: Archive - Source URL: None The way that the operating system decides which processes to run and when is a complicated process known as **scheduling**. This is how it works. We should understand a few vocabulary words before we discuss schedulers in detail. - job: an execution stream for a certain amount of time (not the same as process!) - workload: the jobs that must be executed - latency: time taken for one operation - throughput: number of total operations ## Metrics We can design schedulers that are designed to minimize certain metrics. No scheduler is perfect at everything, so we need to prioritize what metrics matter. - turnaround time: how long does the job take to complete? - \\( completionTime - arrivalTime \\) - response time: how long does the job take to start? e.g. game keystrokes - \\( scheduleTime - arrivalTime \\) - waiting time: how long does the job wait in the ready queue? - same as response except when killing procs - throughput: jobs completed per unit time - resource utilization: manage small resources well e.g. battery - overhead: how strenuous is the scheduler itself? - fairness: how well is CPU time shared between jobs? ## FIFO FIFO stands for **First In, First Out**, which is a fairly self-explanatory title. A FIFO scheduler does jobs in the order that they arrive. This kind of scheduler is trivial to implement. However, we can run into issues if say, the first job takes a long time to complete. The other jobs that would complete much quicker are waiting even though it would be better if we could just get those out of the way first. ## SJF SJF stands for **Shortest Job First**, which again is self-explanatory. This is better than FIFO, but relies on the OS having oracle-like knowledge of all the jobs that will come, since a shorter job could come later that the scheduler can't factor into its calculations. These schedulers only work when you have access to all the jobs at once, which is almost never the case in real-world workloads. We need to use some kind of *preemptive scheduling* which will switch and split jobs as necessary. ## STCF STCF stands for **Shortest To Completion First**. It is similar to SJF, but it recalculates the job closest to completion every time a new job arrives. This way, if a super short job arrives while a long job is executing, STCF can switch to it and complete it quickly before finishing the long job, which improves completion time. This expands the machinery needed in the scheduler but has much better outcomes. ## RR RR stands for **Round Robin**. This approach is fairly different than the the previous two schedulers because it does not make any attempt to optimize for the shortest time. Instead, RR optimizes for fairness by executing every single job on the same exact time intervals. So it will run 10 ms of Job A, then 10 ms of Job B, then 10 ms of Job C, etc. in the fairest way possible regardless of the actual length of those jobs. Preemptive scheduling is pretty cool, but we're still treating all our jobs like they're the same. If we introduce the notion of **priority** in jobs then we can use it for real-time jobs. ## MLFQ MLFQ stands for **Multi-Level Feedback Queue**. The implementation is similar to RR but it works on multiple priority levels. There are two job types: *interactive* and *batch*. Interactive processes are those like games and text editors that need immediate response and therefore low response time. Batch processes are those like daemons that are lower priority and care more about general turnaround time than response time. For our implementation, we have multiple priority queues: ```java if a.priority > b.priority { a.run(); } else if a.priority == b.priority { rr(a, b); } ``` MLFQ prioritizes *nice* processes. This is a technical term that means that a process is satistfied with getting a response and is willing to turn over the CPU as needed without the OS needing to force it to. All jobs begin by having top priority, but if a job takes too long on the RR, then you demote it to a lower priority level. This way, smaller jobs that are contained within RR are executed quickly at top priority. However, this system is incredibly easy to game. Processes control the jobs that they send to the CPU so you can split up jobs exactly aligned to RR to execute all the jobs at top priority even though you don't actually need it. ## Lottery The lottery system is fairer, just like a real lottery. Every process gets a certain amount of lottery tickets associated with it, scaling with priority. Whichever process wins the lottery gets to run its job. This way, higher priority jobs have a higher probability of running than lower priority jobs, but it's protected against gaming the system. The number of tickets that a process gets represents its CPU usage. ```c int counter = 0; int winner = getrandom(0, total_tickets); node_t *curr = head; while curr { counter += curr->tickets; if (counter > winner) { break; } curr = curr->next; } ``` ## Multiprocessing Every scheduler we've discussed so far assumes only a single CPU can execute jobs, which was true for a long time. Now, however, multicore CPUs are becoming increasingly common in consumer electronics. How do we schedule jobs among different CPUs? *Cache affinity* and *coherence* can become issues. CPUs have a cache located next to them different from main memory where commonly used memory is stored in order to speed up operations. Every core has its own caches, so what happens when multiple cores execute the same stream? They have to have cache coherence so that we don't end up with bugs. The way to fix this is to copy caches between cores. Obviously, this is pretty inefficient. We want to avoid this as much as possible by having each process run on its own CPU thread as much as possible; this is cache affinity. Basic FIFO scheduling is simple but is not good at preserving affinity. Having a scheduler that preserves affinity at the cost of some other inefficiencies can actually cause massive speedups because of the importance of cache affinity. However, this machinery can be complex. ## CFS The **Completely Fair Scheduler** is the scheduler that Linux actually uses. Instead of using time-slices to manage job usage, the scheduler assigns processes a proportion of the CPU. The fairness is absolute because each thread gets an absolute amount of CPU running. This also fixes multiprocessing issues because the scheduler is based around CPU cores instead of time. However, the switching rate might change a lot because we are no longer looking at time, and as we discussed context switching has its own costs that we might want to minimize. This is a tradeoff made for fairness. I mentioned niceness earlier. CFS uses a more complex version of niceness which ranges from -20 to 19, with the default of 0, and higher values being worse. There is also a separate priority called *real-time priority* ranging from 0 - 99 which always execute before nice processes, regardless of how nice they are. In order to calculate RR, we take the target latency e.g. 20 ms and split the time across the tasks equally among process of the same priority. However, because this is based around CPU and not constant time, this can be extremely fast. A tree data structure is used for priority in order to quickly get the highest priority and least time jobs. **Target Latency** is the minimum amount of time required to get a task at least one turn on the processor. Within this window, every process gets some CPU. **Minimum Granularity** imposes small unfairness in CFS. There is a floor on the timeslice of 1 ms regardless of how much CPU we have. This means that for a large amount of jobs, smaller jobs get unfairly good treatment. ## Lists in OCaml - URL: https://sharifhsn.dev/blog/lists/ - Structured data: https://sharifhsn.dev/api/posts/lists.json - Description: Lists are the most basic data structure in OCaml. The most analoguous structure is a vector in C++ or Rust. Lists are homogenous and of arbitrary length. - Date: 2022-01-27 - Exact published timestamp: 2022-01-27 - Topics: Principles of Programming Languages - Categories: Principles of Programming Languages - Source: Archive - Source URL: None Lists are the most basic data structure in OCaml. The most analoguous structure is a vector in C++ or Rust. Lists are *homogenous* and of *arbitrary length*. The most basic list is **nil**, the empty list: `[]`. We can prepend elements to a list through the **cons** operator: `::`. Every display of a list is actually just syntactic sugar for every element being cons with the empty list: ```ocaml [1, 2, 3] = 1 :: 2 :: 3 :: [] ``` Importantly, lists are immutable! When they are created, they cannot be changed. If you want to add an element to a list, you must construct a new list based on the old one. > This might seem inefficient, to create a new list every time we want to change it. But the semantics of a language do not necessarily correspond to its compiled execution. In particular, the invariant that lists are immutable can lead to optimizations where a compiler doesn't need to keep track of mutated information. It is convention to call the right hand side the *tail* and the left hand side the *head*. The cons operator has an element on the left side and an list on the right side. Let's think of lists in terms of expressions. `[]` is a value like 0 or 1. The cons operator evaluates the left expression to an element and the right expression to a list of type element. Like with function arguments, the order of this evaluation is irrelevant as long as we don't introduce side effects. Lists can also have the **polymorphic** type `'a`. This is similar to a generic type in languages like Java, C++, and Rust. An `'a` list can have any type. This is most useful when it comes to function types. If the function type accepts or returns a `'a` type, then it is a generic function that can be used on a list of any type. Nil is a `'a` list. You can also have nested lists like this: ```ocaml let m = [[1]; [2; 3]] ;; (* int list list *) ``` Unlike tuples, these lists do not have to be the same size. They do need to have the same type, however, so this is a list of int lists, or an `int list list`. The cons operator *does not* work to concatenate two lists. It can only add one element to an existing list. This is a very common bug (at least for myself). OCaml has a special sugar operand `@` to concatenate lists, although it is technically an ordinary function. ## Pattern Matching **Pattern matching** is an extremely important concept in OCaml for functional programming. It allows for powerful destructuring of enums, collections, and other complex objects. ```ocaml let head l = match l with | (h :: _) -> h ``` This function `head` here matches the list `l` where there is an element `h` that is cons with an arbitrary expression. The `_` indicates that some expression must be present, but we don't care about what expression it is, and in fact we are throwing it away. `h` is also an arbitrary expression, but we are using it as the return type. This match statement is not *exhaustive*. An exhaustive match means that there is an evaluation for every possible case of `l`. Here, there is no case for when `l` is nil. This does not work in OCaml; every match must be exhaustive. Here is an example of an exhaustive match: ```ocaml let rec sum l = match l with | [] -> 0 | h :: t -> h + sum t ``` This `sum` function works for every case of `l`. Either it is nil, or it is a head cons a tail. You will also notice here that the two cases resolve to the same type. This is another requirement for match statements: all patterns must evaluate to the same type, which is not necessarily the type of the initial expression. Here, the initial expression has the type `int list` and the return expression has the type `int`. This function only works for `int list`. This restriction is not present for our `head` function because `h` is polymorphic type `'a`. Pattern matching is generally much better than its alternatives. OCaml will warn you if your function is non-exhaustive, and it will throw an exception for an unhandled case. Also, duplicated cases are easy to avoid becaues OCaml will give a similar warning for unused cases. ## Recursion with Lists In order to manipulate lists in any significant way, we need to use recursion. Lists are immutable, so functions over them must be recursive. ```ocaml let rec length l = match l with | [] -> 0 | (_ :: t) -> 1 + length t ``` We can think of this as recursive because we have our base case of nil and our iterative step of applying the function to successive tails. This function does not use `h`, but we might, as in the `sum` function above. Recursion might seem like it incurs overhead here, as in C-like languages recursion typically costs stack frames. However, OCaml knows that even though you are writing your code recursively, all you're doing is iterating over an immutable list, so it will elide those issues. Such recursive functions are called **tail-recursive** and are essential to the use of functional languages. However, these functions work through the lists forward, like a linked list. How can we do something like reversing a list? ```ocaml let rec rev_aux l acc = match l with | [] -> acc | x :: xs -> rev_aux xs (x :: acc) let rev l = rev_aux l [] ``` The key here is the variable `acc`, which is known as an **accumulator**. This is a variable that is included with each recursive function call that accumulates operations on it. This is a way of simulating side effects in a controlled way which is often useful. `acc` will accumulate heads while the tail grows smaller, and the actual `rev` function will get the accumulated nil that it passed to `rev_aux`. ## Electric Potential - URL: https://sharifhsn.dev/blog/electric-potential/ - Structured data: https://sharifhsn.dev/api/posts/electric-potential.json - Description: Besides electrical charge and field, there's another force that is relevant to electricity: electric potential. - Date: 2022-01-25 - Exact published timestamp: 2022-01-25 - Topics: Physics - Categories: Physics - Source: Archive - Source URL: None Besides electrical charge and field, there's another force that is relevant to electricity: **electric potential**. ## Potential Energy Gravity is similar to the electrostatic force, apart from the fact that it acts on mass instead of electric charge. The equations are practically identical. We're familiar with the idea of gravitational potential energy; why not apply the concept to charge as well? Recall that **work** is expressed as the difference between potential energy. The equation for work done *by* an electric force is: $$W_{AB} = EPE_A - EPE_B$$ where \\(A\\) and \\(B\\) are the positions that the charges are in. ## Electric Potential Difference Remember that the electric force depends on the small test charge \\(q_0\\). If we divide all parts of the previous equation by \\(q_0\\), then we get the work per-unit-charge. For the first time, we're going to be talking about **electric potential**, expressed in **volts** \\((V)\\). This is different from electric potential energy, which is expressed in joules (\\(J)\\). The equation for electric potential derived from EPE is \\(V = \frac{EPE}{q_0}\\), showing that volts are the same as joule/coulomb. \\(V\\) can only truly be measured as \\(ΔV\\), becuase it is relative measurement in terms of work, similar to energy. This potential difference \\(ΔV\\) is known as **voltage**. When we talk about voltage in common parlance, that is the electric potential difference between two charges, typically the terminals of a capacitor. We can think of positives and negatives here like gravity. Because electricity flows from positive to negative, that's the direction of "gravity". Positive charges accelerate to lower electric potential like massive objects accelerate to lower heights. Negative charges oppose this acceleration like the normal force to higher electric potential. You might have also seen the unit \\(eV\\), \\(MeV\\), or \\(GeV\\) with reference to energy. This is the unit *electron-volt*, which is, as implied, the charge of an electron multiplied by one volt. This value is \\(1.60 × 10^{-19} J\\), and the mega and giga versions are six and nine order magnitudes greater than it, respectively. ## Volts and Point Charges If we apply this concept to the point charges we have discussed, we get a pretty simple formula: $$ΔV = \frac{-W_{AB}}{q_0} = \frac{kq}{r_B} - \frac{kq}{r_A}$$ And in order to determine \\(V\\) for the \\(q\\) at just point \\(A\\), we consider point \\(B\\) to be infinitely far away and eliminate it, leaving us with the final potential equation: $$V = \frac{kq}{r}$$ Note here that \\(q\\) is not absolute. \\(V\\) will have the same sign as \\(q\\). Like with electrostatic force, we can calculate \\(V\\) with respect to many point charges by adding them together, minding our signs . We can calculate electric potential energy for multiple point charges in a similar way by adding together every pair EPE \\(\frac{kq_1q_2}{r}\\). ## Equipotential Surfaces An **equipotential surface** is one where the electric potential is the same everywhere. Think of a hollow sphere which surrounds a point charge. Every point on that surface is the same distance from the charge and so will have the same electric potential. If you have a charge on an equipotential surface, then no work is done on it by the net electric force if it moves along it. This is because the potential is the same everywhere, looking at the equations from earlier. On the same equipotential sphere, the electric field is perpendicular to the surface and points to decreasing potential, like gravity. The most clear example of this is a parallel plate capacitor. If we try to calculate the electric field from electric potential, it is simply $$E = -\frac{ΔV}{Δs}$$ where \\(Δs\\) is an expression of displacement, like distance. This is what we would use to find the electric field or distance for a parallel plate capacitor. ## Capacitors and Dielectrics We have already discussed one kind of **capacitor**, the parallel plate capacitor. A capacitor is simply two conductors placed close together, which are not touching. There will often be an insulator placed in between them called a **dielectric**. The capacitor stores electric charge, with the positive plate having more electric potential. That potential difference is the voltage of the capacitor, although the plates have the same charge. **The magnitude \\(q\\) of the charge on each plate of a capacitor is directly proportional to the magnitude \\(V\\) of the potential difference between the plates:** $$q = CV$$ **where \\(C\\) is the capacitance in farad \\((F)\\).** As you can tell from the equation, capitance is coulomb/volt and it represents how well the capacitor can store charge. Like coulombs, it is typically expressed in smaller forms like \\(μF\\) and \\(pF\\) which are six and twelve orders of magnitude less. We can increase capacitance by adding a dielectric in the capacitor. The molecules in the dielectric will form dipole moments oriented towards the positive and negative plate as appropriate. These dipole moments will stop the electric field from passing through the dielectric, and instead some of it will begin and end at the surface. This reduces the strength of the electric field and allows it to store more charge. The equation for capacitance of a parallel plate capacitor based on the dielectric constant \\(κ\\) is: $$C = \frac{κε_0A}{d}$$ There is also work done to fill a capacitor with charge, because it stores energy. The work done by a battery, for example, to fill up a capacitor's charge is expressed in multiple ways as: $$Energy = \frac{1}{2}qV = \frac{1}{2}CV = \frac{q^2}{2C} = \frac{1}{2}(\frac{κε_0A}{d})(Ed)^2$$ ## Functions in OCaml - URL: https://sharifhsn.dev/blog/functions/ - Structured data: https://sharifhsn.dev/api/posts/functions.json - Description: A function takes arguments, performs operations using them, and returns a value. In OCaml, we can write functions like so - Date: 2022-01-25 - Exact published timestamp: 2022-01-25 - Topics: Principles of Programming Languages - Categories: Principles of Programming Languages - Source: Archive - Source URL: None A **function** takes arguments, performs operations using them, and returns a value. In OCaml, we can write functions like so ```ocaml let rec fact n = (* the function is named fact and its arg is n *) if n = 0 then 1 else n * fact (n - 1) ;; ``` This is a *recursive* function, so we must prepend the keyword `rec` to the name of the function. The reason for this is that the function name is not automatically in scope, so if we don't use that we won't necessarily know what `fact` refers to. `rec` adds the function name into scope. As you can see, the declaration of a function is fairly similar to the declaration of a variable. In fact, we can use functions as variables in OCaml! This is because functions in OCaml are *first-class functions*. You also might have noticed here that we haven't mentioned the type of anything. This is because the type of the function can be inferred, also known as **type inference**. This process is a normal part of type checking for correctness. The compiler itself will look for a type that the code is correct for. In this example, `n` is compared with an integer, so it must only be an `int`. The return type is `n` multiplied by its own return type, so that must also be `int`. The compiler can see without us telling it that the function is of type `int -> int`. But what is that `int -> int` that I just wrote? That is the constructor for a function type. The last type indicates the return type, and all other types are the arguments in order. The function type `float -> int -> float`, for example, takes a `float` and `int` as its arguments and returns a `float`. Try looking at different OCaml functions and guessing its type as a fun exercise! :smile: > Aside: you might question why we couldn't just get a float as a result from the multiplication. The standard operands in OCaml only apply to `int`, and applying them to other types is an error. In order to add `float`, we must use `+.`, not `+`. Type checking works like a proof by contradiction. We assume that a function has some type, then check if there are any logical contradictions created by this assumption. If there are, then we reject this type. ## Calling a Function Function calls (also known as *applications*) look very similar to function declarations; they are simply the name of the function followed by its arguments, no parentheses necessary. The type checking involved is that we check if every expression supplied as an argument corresponds to the same type as it's supposed to. We evaluate each expression from right to left to see if they're correct or not (though this doesn't matter unless you have side effects). Once we find the value for each argument, we replace each argument in the function with its value, which creates an expression that can be evaluated to the final value. > Aside: Although we have type inference, we can use *type annotations* with OCaml. These increase the readability of code by making the type of the variable obvious: `let (x : int) = 3` ## Electrostatics - URL: https://sharifhsn.dev/blog/electrostatics/ - Structured data: https://sharifhsn.dev/api/posts/electrostatics.json - Description: What is electricity, and how does it work? - Date: 2022-01-18 - Exact published timestamp: 2022-01-18 - Topics: Physics - Categories: Physics - Source: Archive - Source URL: None What is electricity, and how does it work? ## Electrons Fundamentally, the negative charge comes from the **electron**, a subatomic particle. The positive charge comes from **protons**, also in the atom. The unit of the charge is the **Coulomb** (\\(C\\)). An electron has a charge of \\(-1.60×10^{-16} C\\), and a proton has the same charge but positive. However, the proton is much heavier than an electron, being about 1836 times as massive. This means that electrons are easy to displace, which is why their movement is the fundamental concept behind electricity. The way that electric charge is generated is by rubbing two materials together in a way that transfers electrons between the two. For example, if you rub your clothes against a shag carpet, it will transfer electrons to your clothes. *However, charge is neither created nor destroyed, due to the law of conservation of electric charge!* As you might already know from playing around with magnets, like charges repel and unlike charges attract each other. This attraction is known as the **electrostatic force**. ## Conductors/Insulators Electric charge is not limited to its power to attract and repel. It can also move through objects; this is how electricity travels from power plants to your house, through wires. Not all materials are created equal in their ability to **conduct** electricity through them. Those that conduct poorly are called electrical **insulators** and those that can well are called electrical **conductors**. This material property is typically related to how easily valence electrons can move around in the molecules of the material. ## Charging by Contact/Induction I said before that electrical charge can be generated by rubbing, for example against a shag carpet. This is known as **charging by contact**. However, this is not the only way to charge an object. As we know, electrons can be attracted and repelled. If a negatively charged object is brought close to a neutrally charged object, the electrons in the neutrally charged object will be repelled to the other side, changing the distribution of charge in the object. This is known as **charging by induction** because the electrons are induced by the charge to move. This effect will disappear when the inducing object leaves, unless we introduce **grounding**. A grounding wire is a way to disappear free electrons from a charged object. Typically, an object is wired to a large diffuse object such as the earth, so free electrons will leave the small object forever. In the example I used above, if there is a grounding wire attached to the other side where the electrons were repelled to, then those electrons would "disappear". After the inducing object leaves, the formerly neutral object would become positively charged. This can even occur in materials like plastic which are insulating. Although the free electrons cannot move as in materials like metals, positive charges can still be induced towards the surface, allowing for a surface positive charge that negatively charged objects to stick to. ## Coulomb's Law **The magnitude \\(F\\) of the electrostatic force exerted by one point charge \\(q_1\\) on another point charge \\(q_2\\) is directly proportional to the magnitudes \\(|q_1|\\) and \\(|q_2|\\) of the charges and inversely proportional to the square of the distance \\(r\\) between them:** $$ F = k\frac{|q_1||q_2|}{r^2} $$ **where \\(k\\) is a proportionality constant: \\(k = 8.99 × 10^9 N \cdot m^2/C^2\\) in SI units.** So what does this law mean? Crucially, this law gives no mention of direction. The magnitude of the force is the same regardless of whether the charges attract or repel, so we don't need to worry about that. > The constant \\(k\\) is sometimes written as \\(k = 1/(4πε_0)\\) where \\(ε_0\\) is the *permittivity of free space* with the value \\(8.85 × 10^{-12} C^2/N \cdot m^2\\). A Coulomb is an absolutely massive amount of charge that is typically only found in situations like lightning. It is far more common to encounter microcoulombs expressed as \\(μC\\) which are \\(10^{-6} C\\). This is a situation that is expressed between only two point charges. What happens when we have multiple point charges? We have to calculate the *vector sum* of all these forces. This is easiest when all the charges are in a line, as this is simply a matter of adding and subtracting the forces depending on whether they are attractive or negative. We must calculate Coulomb's law for each pair and then combine those as appropriate. When we have points on a plane, things get a little more complicated. The easiest way to handle this is by separating each force into its \\(x\\) and \\(y\\) components. Pick one point to be set at the angle 0° degrees with respect to the point we are calculating for, then set the \\(θ\\) for each other point with respect to that line. Then, get \\(F_x\\) and \\(F_y\\) through \\(\cos{θ}\\) and \\(\sin{θ}\\), respectively. The point with 0° will only have an \\(x\\) component. Then, once you have the \\(F_x\\) and \\(F_y\\) all parcelled out according to positives and negatives, get the total \\(F\\) by applying the Pythagorean theorem to them. The answer \\(θ\\) is given by \\(\tan^{-1}{\frac{F_y}{F_x}}\\). ## The Electric Field We have so far been talking about how individual point charges affect other point charges. When we talk about an entire system of charges, we need a broader, high-level understanding of charge. For that purpose, **the electric field** is essential. Fundamentally, the electric field is defined as the electric force per coulomb \\(\vec{E} = \vec{F}/q_0\\). This calculation pretends that we have a small positive *test charge* \\(q_0\\) placed in the middle of this electric field, which is affected by the various charges to have a force pointing a direction. This test charge typically has a small magnitude so as not to disturb the surrounding field. The electric field is independent of whatever point charge it's affecting. We could imagine a field with a magnitude \\(2.0 N/C\\) and direction upwards that causes a positive charge to go in the same direction as it, but a negative charge to go in the opposite direction. We can calculate the magnitude for each point charge with \\(F = |q_0|E\\). Electric fields, just like forces, are vectors. When we add fields together, we must equally be mindful of \\(E_x\\) and \\(E_y\\). An electric field is composed of multiple point charges. The amount of electric field that an individual point charge \\(q\\) produces is expressed by the equation $$E = \frac{k|q|}{r^2}$$ Notice how \\(q_0\\) the test charge is not included in this equation. Since this new equation only expresses magnitude, we must individually determine direction by the sign of the charge. If \\(q\\) is positive, \\(\vec{E}\\) is directed in the opposite direction of \\(q\\), and vice versa for negative. An additionally useful equation relates to **parallel plate capacitors**. These are two parallel plates of opposite charge that have an electric field directed from the positive plate to the negative plate. The equation for this electric field is $$E = \frac{q}{ε_0A} = \frac{σ}{ε_0}$$ As you can see in this equation, the symbol \\(σ\\) denotes \\(q/A\\) and is sometimes called *charge density*. Importantly, the distance between the plates is actually totally irrelevant to the electric field, unlike with point charges. ## Electric Field Lines It is often useful to draw fields as a set of lines directed out of or towards a point charge. ***In general, electric field lines are always directed away from positive charges and toward negative charges.*** For convenience, we usually only draw them in two dimensions, though technically charges act in three dimensions. The number of lines is expressed as proportional the magnitude of the charge, so five times as many lines for five times the magnitude. In general, the electric field is stronger when lines are closer together, which is why parallel plate capacitors have the same electric field in the middle but bulge out at the ends. The lines are not always straight, as in the case of an **electric dipole**. An electric dipole has two separated point charges of opposite signs and equal magnitude. The field lines in this case come from the positive charges and curve towards the negative charge, getting faster as they get closer. The electric field vector in this case is *tangent* to the electric field at any point. Between two charges of the same sign, there is practically no electric field because the lines curve away from each other. ## Shielding Electric fields can exist anywhere where electrons can move freely, which includes the inside of a conductor. It's a logical conclusion: electrons hate being near each other, so they try to get away in the conductor. That leaves them all at the surface of the conductor. The same principle applies for positive charges; conductors generally have all their excess charge at their surface. However, since all the excess charge is at the surface, that means that in the interior of the conductor, there is no movement of free electrons because they're not there. This means that there is *zero* electric field inside a conductor. If we imagine a conductor within a parallel plate capacitor, there is a lot of electric field going on. However, all of the charges are directed outside of it. Inside the conductor, any charge is completely shielded from the intense electric field outside the conductor. ## Gauss' Law We have so far discussed electric fields created by point charges. In fact, an electric field is often composed by multiple charges that are spread out, also known as a *charge distribution*. Understanding an electric field created by a charge distribution is significantly more complex and requires understanding **electric flux**. We can think of flux on a high level as similar to a vector, with a direction and magnitude, but with the additional dimension of area. This concept will be important when discussing magnets. We can consider the formula for a point charge as a special form of **Gauss' Law** which describes electric fields in general. Remember how we can substitute \\(1/(4πε_0)\\) for \\(k\\)? This means that our new equation is \\(E = q/4(πε_0r^2)\\). If we multiply both sides by \\(A\\) which is \\(4πr^2\\) for a sphere, then we get $$EA = \frac{q}{ε_0}$$ where \\(EA\\) is electric flux, also known as \\(Φ_E\\). However, we don't always have a nice, neat distribution of charge like in a sphere. A Gaussian surface can be arbitrarily shaped. We have to think of this similarly to how we think about integrals in calculus. If we take an arbitrarily small area \\(ΔA\\), it has an electric field that is essentially constant in magnitude and direction. The surface has a *normal* which is a line perpendicular to the surface. The angle \\(ϕ\\) is between the electric field vector and the normal. With all these variables, Gauss' Law is expressed as so: **The electric flux \\(Φ_E\\) through a Gaussian surface is equal to the net charge \\(Q\\) enclosed by the surface divided by \\(ε_0\\), the permittivity of free space:** $$Σ(E \cos{ϕ})ΔA = \frac{Q}{ε_0}$$ **with the left side being \\(Φ_E\\) in \\(N \cdot m^2/C\\).** The flux has the same sign as \\(Q\\). ## CPU Virtualization - URL: https://sharifhsn.dev/blog/cpu-virtualization/ - Structured data: https://sharifhsn.dev/api/posts/cpu-virtualization.json - Description: We need a way to map on to the physical CPU through our operating system. We do this through virtualization. - Date: 2022-01-04 - Exact published timestamp: 2022-01-04 - Topics: Operating Systems Design - Categories: Operating Systems Design - Source: Archive - Source URL: None We need a way to map on to the physical CPU through our operating system. We do this through virtualization. ## Processes A **process** is an *execution stream* in the context of a process state. It is a self-contained stream of executing instructions on the CPU. Each process has a state which is composed by everything that the code can affect or be affected by, for example *registers*, address space, open files, etc. We need this state so that we can pause and resume processes without resetting the whole thing, just by picking up where the state left off. You can think of address space as the region of memory in which the CPU operates, though this is a simplification due to memory virtualization. A process is technically different than a program. When we talk about a program, it's usually a collection of files on our computer. The process refers to the actually executing code which is dynamic with respect to its code. We can have multiple processes executed that are the same program. Processes do not share information with other processes. This is distinct from the idea of threads. If a process examines the "same" memory address with respect to its address space as another process, they will see different values. Even though both process are looking at `0xFFE84264` the memory virtualization means that they are looking at different physical memory. Threads share data so they have the same address space; they are in a way a lightweight process. ## Address Space As mentioned, every process has its own address space. The OS assigns a chunk of memory towards it (again, this memory is virtual). At the high address, there is a stack of local variables that grows downwards. This stack contains the execution of the program, such as functions being called. The code segment which is the actual program is read from the absolute bottom. There are a few segments like `.data` and `.bss` which refer to global initialized/uninitialized variables, respectively. Then there is the heap, which is dynamic and grows upward. Usually, the heap is much larger than the stack because it contains allocated memory for the process. > The stack and the heap grow towards each other, so you might think they would collide at some point. In reality, there is a wide enough chasm between them that if they collide, you have bigger problems to worry about. ## CPU Virtualization Our problem is that a CPU can only execute one execution stream at one time. We want processes to be running at the same time, so how do we achieve the illusion that the process has full control of the CPU? One solution is *direct execution*. We will run the process as simply as possible and execute all of its instructions in sequence. This is extremely simple to implement, but it means we can only run one process at once. If the process runs forever, then the CPU is stuck. Processes might write to other data that it's not supposed to, or do some slow I/O operation, or execute privileged instructions it's not allowed to accesss. In general, we don't want to trust processes to behave. Operating systems use *limited direct execution*, which maintains some control over the execution of the process instead of giving it unfettered access to the CPU. ## System Calls We want to make sure user process can't harm other processes. CPU hardware supports privilege levels for security reasons. These instructions should only be run in kernel space by the operating system, and user processes in user space can not access them. Privileges can have multiple levels. In order for processes to access privileges, we use **system calls** that can either pass the function off to the OS or give the process privilege, also known as a *trap*. ```c ID1 = syscall(SYS_getpid); // syscall ID2 = getpid(); // call to libc ``` Let's say a process wants to execute `read()`, which is a syscall. It will move the syscall ID for read `0x6` into the `%eax` register for execution, then run the instruction `syscall`. The operating system will then read the trap table to figure out what to do. The trap table is located in the hardware, which typically has an entry for system calls. Other traps could be `illegal access` for memory region that it doesn't have access to. A trap just refers to a hardware operation that triggers on some software interaction. This is why Javascript can't just randomly hack your computer from the internet; it is executing within a user process e.g. Chrome and is limited by the hardware traps in what it can do. ## Multiprogramming We want to make sure that multiple processes can run at once. This means we must switch between processes. There are two components to this; how to switch, and when to switch. The dispatch loop has a simple structure, but the devil is in the details. ```c while (1) { // run process A for some time-slice // stop process A and save its context // CONTEXT SWITCH // load context of another process B } ``` We have multiple ways to context switch. One way is *cooperative multi-tasking*. We trust the process to give up the CPU when it has judged some amount of work has been done or time has passed. We will provide `yield()` syscall for magnanimously donating CPU time to another process. However, this is annoying to program in, so most programs do not do this. Modern programs use *true multi-tasking*, which gives OS control over which processes are running. The hardware will generate a timer interrupt every 10 ms, for example (Linux) that the OS will decide when to switch or not. ## PCB The **process control block** is a descriptor of a process that saves its context into a `struct`. There is a `struct` for every process that is running in the CPU. A PCB stores the following information: - PID - unique identification for a process - Process state (`enum` running, ready, or blocked) - Execution state (registers at time of pause) - Scheduling priority - Accounting information (parent/child procs) - Credentials (resource access/owner) - File pointers When we save the context of the process in the dispatch loop, the PCB is where it is saved. It is typically less than a few kilobytes depending on the complexity of the process. The PCB is stored on the *kernel stack* which is a stack data structure located in kernel space. Switching a context means moving the stack pointer to the process that is going to resume and then retrieving that PCB from the stack. Then, we exit kernel space and resume executing B in user space by restoring the PCB information to the CPU. When a process is doing I/O, it's not using the CPU. This is an excellent point where the OS can block the process and run some other processes that actually need the CPU. Once the I/O is done, it is moved from the blocked state to the ready state. The convention is that only ready state programs can go to the running state. ## Process Creation One way to create a process is just creating one from scratch. We will load the code we're running into memory at the bottom of the stack and create the stack. We also create a PCB with a kernel stack and everything, then put the process in the ready state. However, there are lots of complicated parts of processes that are difficult to initialize from scratch e.g. permissions, I/O, environment vars so this is generally not preferred. A better way is to clone an existing process and change it to the appropriate information. `fork()` clones the calling process and `exec(char *file)` replaces the current process with the new process. When we fork, we save the current state of the program as a new PCB that is added to the kernel stack, which is very optimized to use copy-on-write semantics. > The basic idea of copy-on-write is that it uses a pointer/reference for memory until we need to modify memory. This is an incremental process that trades off some overhead for a pay-as-you-go memory saving system.