Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Python For Defenders

Python for Defenders: Code, Detect, Defend.

License/Copyright

Although this repository is open source and suggestions in the form of Pull Requests are welcome, this remains the intellectual property of The Taggart Institute, LLC, under the following license:

Creative Commons License
This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

AI-Free Disclaimer

No part of this book was generated by a large language model such as ChatGPT, Claude, or Gemini. The prose and code you see here was created by humans, mostly by me, Michael Taggart, with help from open source software authors and contributors.

How to Use This Repo

This repository is intended for use in conjunction with the course Python For Defenders.

To use the repo,

git clone https://github.com/The-Taggart-Institute/python-for-defenders
cd python-for-defenders

Changelog

  • 2026-10-07: Complete rewrite and unification of parts 1 and 2
  • 2026-03-12:: Update dependencies
  • 2022-09-21:: Added Usage; fixed test for 1-2.

0-1: Welcome

You might be wondering why folks on the defensive side of cybersecurity need a specialized course on the Python programming language. Unfortunately, even today Python programming is not a common skill set amongst most defenders. This lack of capability means that many defensive teams are not as efficient or effective as they could be, and so I would like to share my knowledge as a Python developer and a defender to address that gap.

Learning Objectives

At TTI, we split our learning goals into Skills—things you should be able to do, and Concepts—things you should be able to understand.

Skills

By the end of this course, you should be able to:

  • Write Python scripts and Jupyter Notebooks
  • Manage Python project dependencies
  • Use Python to parse and analyze data sources
  • Use data science tools techniques for analysis
  • Interact with remote APIs via Python
  • Create dynamic reports with Python

Concepts

By the end of this course, you should understand:

  • Python syntax
  • Jupyter notebook execution
  • The value of literate programming
  • The tool creation process in Python
  • When and why to build Python tools
  • How to craft useful, reusable notebooks for use in common processes

Course Structure

Previous versions of this course were split into two parts. Part 1 was focused on Python basics, and Part 2 applied those skills to defensive operations. These two components have since been combined into one complete course of study.

Most of the course text is within the course’s Jupyter notebooks. This book will provide additional context as well as the original video lessons when appropriate.

Prerequisites

This course does not assume familiarity with programming, but if you’ve had any exposure to writing code at all, it will be helpful. You will require familiarity with the Linux command line. If you need to start there, we have you covered.

Familiarity with basic networking concepts (IP addresses, subnets, etc.) will also be valuable.

Required Materials

You’ll be running Jupyter Notebooks on your local computer. That means you’ll need a way to install Python and Python packages. This course uses uv for managing Python versions and dependencies. For Windows users, I strongly recommend using Windows Subsystem for Linux as your base platform. Mac/Linux users, you can use uv natively or in a guest VM, but that’s on you to set up.

Previous versions of the course referenced an official virtual machine. Maintenance of that VM was too burdensome and not valuable enough to continue. Between WSL and macOS being macOS, you should have what you need. Linux users, odds are you have Python installed anyway.

About AI

“Why bother learning Python, when AI can do all of this?”

While large language models are great at creating Python syntax, they do not understand anything. It is still on you, the defender, to grasp the nature of a problem and design solutions. What’s more, only experts in a subject area can identify when AI goes off the rails. The only way to gain true expertise in an area is doing it yourself. Letting the model write the code is faster (sometimes), but we’re not interested in speed at the moment. We’re interested in learning. If you want to truly understand this topic, I beg you: put the agents aside for this course. We will proceed for the most part as though LLMs do not exist, and that we must (and should) rely on our skills as defenders and programmers.

With that, again, I want to welcome you to Python for Defenders. Up next, we’re going to talk about why Jupyter notebooks are so cool and why we’re using them in this context.

0-2: Do I Need to Learn to Code?

0-3: Why Notebooks?

There are a couple questions worth answering before we get underway.

  1. Do security analysts even need to learn to code?
  2. Why Jupyter Notebooks as a vehicle for introducing Python?

Do Analysts Need to Code?

Unequivocally, no they don’t. One can have a perfectly successful career in cybersecurity without learning to program—especially in the age of LLMs (I said we’d mostly ignore them). However, adding this skill set creates a lot of possibilities for a defense team. Automation, data parsing and transformation, deep analysis, and custom tooling are some examples of why you want some programming ability amongst the analysts. You can even make the tools you already have more effective by building the missing integrations or middleware to make them interoperate. After all, you’re paying for tools that have APIs, so you might as well take advantage of them.

Why Jupyter?

For this one, we’re going to rely on Donald Knuth, the grandfather of the craft who wrote the book The Art of Computer Programming. He literally wrote the book!

In another of his books, Literate Programming, Knuth lays out a paradigm for writing software that centers other programmers, not computers, as the “audience” of written code.

In the book’s introductino, he writes:

Let us change our traditional attitude to the construction of programs. Instead of imagining that our main task is to instruct a computer what to do, let us concentrate rather on explaining to human beings what we want a computer to do.

Of literate programmers, he writes:

The practitioner of literate programming can be regarded as an essayist whose main concern is with exposition and excellence of style. Such an author, with thesaurus in hand, chooses the names of variables carefully and explains what each variable means. He or she strives for a program that is comprehensible because its concepts have been introduced in an order that is best for human understanding, using a mixture of formal and informal methods that reinforce each other.

As we’ll see, Jupyter’s structure singularly enables literate programming. By naturally embedding “cells” of executable code with rich text, a Jupyter notebook indeed puts the human audience for the code at the forefront.

For defenders, the resulting artifact can become runbook, automation, and documentation all in one. Repeated over and over again, these notebooks can become a core part of security operations—infinitely flexible, extensible, and customizable for your needs, and as well-documented as your team’s ability to write them.

I think this will make more sense as we set up the working environment. Let’s go do that.

0-4: Environment Setup

Let’s prepare our environment for Python programming. Python is cross-platform, and so are our setup instructions. You can perform this setup on Linux, macOS, or Windows.Even BSD will be fine! However, I strongly encourage a Unix-like operating system as the base. As mentioned in the Intro, I would use Windows Subsystem for Linux if you’re on a Windows OS to easily get started in a Linux environment. Regardless of your OS choice, our first setup task is to install uv.

…after I explain why we need to.

The State of Python in 2026

The Python language is a vibrant, thriving ecosystem. The fields of data science and machine learning adopting Python as the language of choice has even further cemented its ubiquity. But Python has a messy side as well. Managing Python dependencies and managing versions can get complicated quickly. The pip package manager, native to the Python ecosystem, has been disfavored by many Linux distributions which have opted to provide Python packages with their native repositories via apt or dnf. The Python Project format, pyproject.toml, has no official tool to manage it or its declared dependencies.

As a result, many tools have emerged to solve these issues. Previous iterations of this course have used Poetry for dependency management. Poetry is still a valuable tool, but it doesn’t solve all the problems of managing Python. We want as simple a solution as possible for our tooling so we can focus on writing code. That’s where uv comes in.

uv

Why uv?

The uv tool from Astral is a Python distribution, package, and project manager all in one. Not only will it handle project dependencies, but it will also install multiple versions of Python itself. It can also handle installing standalone Python binaries you might want to use on your system.

We will use uv for our Python tooling in this course. If you’ve done any work with Python in the past, this might take some getting used to, as some familiar commands will be preceded by uv.

Installation

Head over to the documentation site for uv. Keep these docs handy; they’ll help you get familiar with the uvtool syntax.

On the installation page, follow whichever installation method you feel most comfortable with. I recommend the standalone installer, although I encourage the curious to read through the installer script.

Once the installer is finished, confirm the command is available on your shell by running uv self version. If it isn’t, you may have to restart your shell session (close the terminal app/reopen it, among other methods).

Install Python

Now that uv’s installed, we can use it to install multiple versions of Python. See what’s available with uv python list.

The list you see will contain <download available> listings, and likely also some discovered Python versions on your system. At the top of the list, you’ll see some with rc in the version. These are release candidates that have not yet reached general availability. I recommend installing the most recent major version that’s generally available. As of this writing, Python 3.14 is the most recent major version. Therefore, I’d run:

uv python install 3.14

You now have that version of Python globally installed and available.

This Repository

In whatever directory you prefer, clone this repository and enter it.

git clone https://codeberg.org/The-Taggart-Institute/python-for-defenders
cd python-for-defenders

Initialize Virtual Environment

While you can—and often do—install Python packages globally on your system, often you want to avoid that for project-specific work. What if Project A uses version 4.2 of a package, while another requires 5.0? To isolate package dependencies—and even Python versions—we use virtual environments for per-project isolation. We can intialize a new virtual environment in our repository with:

uv venv

You’ll be told you can activate with a source command. Do that now. For Bash shells, that will look like:

source .venv/bin/activate

For fish shell users (like myself), there’s a .venv/bin/activate.fish. There’s also .ps1 for PowerShell users.

That’ll change your shell a bit! Your shell has been modified to consider this .venv as your local Python environment. Prove it to yourself. If you’re on Bash, you can try which python to confirm that the python command resolves to a file in this .venv directory. For PowerShell users, this would be Get-Command.

I won’t keep providing PowerShell alternatives. I already suggested WSL, but if you’ve chosen the path of pain, I cannot help you.

You don’t always need the virtual environment activated. To deactivate it and return to a normal shell, simply run deactivate.

A handy feature of uv is that it will look for local .venv folders and use that as its context for certain commands. In fact, let’s run one right now.

Install Dependencies

Take a look at the pyproject.toml file in the root of the repository. You’ll see some metadata about the project, but more importantly, you’ll see a list of package dependencies. You’ll see other Python projects list these in a file called requirements.txt, which is meant for the pip package installer. That’s fine, but pyproject.toml is the new standard. It does more than just list dependencies, but we’ll explore those functions as they come up.

For now, we’ll use uv and this file to install dependencies. Run the following command, either with the venv activated or not.

uv sync

This “syncs” the state of the project to match what pyproject.toml declares. You should see quite a lot of package installation occurring. But where did they go?

Inside of the .venv directory is a lib subdirectory that contains libraries for all installed Python versions in the environment. You can explore this structure to find what you just installed. You might also want to run du -sh ./ to see just how much space this takes.

Running Commands

Our environment is now set up for use, but I want to make sure you understand how to run commands in this environment. Essentially, you have two choices: activate the venv, or uv run. We’ve already covered how to activate the venv in a given shell, but uv run will execute commands as though the venv were activated, even if it isn’t. Give this a try (with the venv deactivated):

uv run which python

You should see the path to the venv version of the Python binary, instead of a global installation like /usr/bin/python.

This also means you can run Jupyter directly without entering the venv.

Launching Jupyter

This is it! Once Jupyter Lab is up and running, we’re reading to get started. With the venv activated, run:

jupyter lab

With the venv deactivated, run:

uv run jupyter lab

Either way, a browser window should launch with Jupyter Lab running.

If a browser window doesn’t open, copy the URL from the terminal (including the token!) to a browser.

To stop Jupyter, hit Ctrl+C.

That’s it! Let’s get started on Jupyter Basics.

1-1: Cells and Variables

Welcome to Jupyter! We’re going to get familiar with the interface, then cover the basic unit of computing in Jupyter: cells.

Interface Tour

The Launcher

When you open Jupyter, you’ll see a tabbed and paned interface with a Launcher front-and-center.

This launcher is the easiest way to kick off a new Jupyter Notebook, but it also has launcher for a built-in Terminal, raw text editor, and even a Python editor with syntax awareness. Mostly, we’ll use the Notebook button, but the others can be handy. You can always reopen this launcher from File → New Launhcer in the menu bar. All the options are also available in the File → New submenu.

Browser

On the left of the interface, you’ll see a file browser.

This allows you to quickly navigate through the directory structure under which Jupyter was launched. Note that at the top of the browser, that directory is noted as /, or the root. Jupyter won’t let you navigate outside that directory.

Note

Jupyter will follow symbolic links, so don’t do anything goofy like put a link to a sensitive directory inside wherever Jupyter runs.

Right-clicking on files and folders in the browser will bring up common file operations. You can also right-click in the empty space to create new files, folders, Python-specific files, and Notebooks.

If the file browser is taking up too much screen real estate, you can collapse it by clicking on the folder icon on the top left.

Cells

Let’s start using this thing. With the Launcher or any other method, create a new Notebook. It can be named Untitled.ipynb for now. This is just a demo Notebook.

Take a look at the tab and the buttons just below the tab title.

Some important points of interest here.

  1. That dot indicates the file is unsaved. Jupyter doesn’t save files automatically!
  2. The save button. Ctrl+S also works. Children, it is a floppy disk.
  3. The “Play” button executes the current cell. More on this shortly. The stop button stops execution of the current cell.
  4. This dropdown determines what kind of cell you’re writing. The three options are Code, Markdown, or Raw, but we’ll only ever use Code or Markdown.

Making Cells

Let’s create a Markdown cell to start. Change “Code” to “Markdown” in the dropdown, and in the cell, enter some Markdown text like:

# This is Markdown

Here's some markdown text.

Then, either press the play button or Shift+Enter to execute the cell and create a new one underneath it. You’ll notice the cell type reverts immediately to Code. So let’s write some code.

print("Hello, Jupyter!")

Again use the Play button or Shift+Enter to execute the cell. A few things happen when you do.

  1. Jupyter is telling you the order of execution of cells. This is the first cell executed in the notebook. Running it again will change this number. This is handy if you have later cells that depend on previous cells that need to be re-run due to updates.
  2. The cell has output! Whenever we ask our Python to generate output, either with a print() statement, invoking an expression, or other output, it will appear beneath the cell. You can right-click on the output and choose “Clear Cell Output” to reset the cell. You can clear outputs of an entire Notebook, resetting it to initial state, with Edit → Clear Outputs of All Cells. You can also fully reset the notebook’s memory with Kernel → Restart Kernel and Clear Outputs of All Cells.

But what is a kernel anyway?

The Kernel

In Jupyter, a “kernel” is both a programming language interpretation/execution layer and the term for a running process for a given notebook. We’re using ipykernel for Python, but there are many others. You can even use Jupyter with PowerShell!

I need to stop mentioning PowerShell.

The kernel maintains state for the Notebook. We’ll find out more about that when we get to variables, but if you want to completely start a notebook’s execution from scratch, you want to restart the kernel.

Cell Output

Let’s fill in a second cell with a simple expression.

5+5

And again, run the cell. You’ll see output beneath the code. This is important. Jupyter will display the result of the last expression in a cell as the output. If you add 2*2 on another line in the same cell and re-run it, the output changes from 10 to 4. If you wanted both results to display, you’d need to wrap both in a print() call, like so:

print(5+5)
print(2*2)

Variables

In Python, like all programming languages, we can store values for later use under named buckets called variables. Different values have different data types, as we’ll soon see. But I want you to be familiar with the idea of variables right away, and that cells share variable values.

Run the following expression in a new cell:

x = 5

Then in another new cell, we simply ask for x

x
# => 5

Two important discoveries here. First, we assign a value to a variable with the = operator. A single equals sign. It might seem weird that I’m specifying one, but it will matter later. Second, we’ve demonstrated that variable values persist across cells. However, execution order matters. I can’t declare a variable in a cell that executes after one in which I reference it.

But if I decide I need a variable, I can make the cell and then reorder the cells in my notebook.

Variables are one of the 4 Fundamentals of Programming. We’ll see more of them as we go on. These are language components that all (or nearly all) programming languages share.

Type Hints

In this course, you’ll see type hint notation like:

message: str = "Hello, World!"

The : str component is the type hint. As we just saw with x = 5 above, they’re not required. However, they are good practice in Python. Not only do type hints tell other programmers what data type you intend a variable to be, but many developer tools like linters (automated processes that review your code) will detect violations of type hints. Python itself won’t, but the linter will.

“Wait, why won’t Python catch it?”

Python is a dynamically typed language, meaning that variables can change times. Try this a cell:

x = 5
x = "five"
x
# => 'five'

Python doesn’t care at all that you changed data types. Try it again with type hints!

x: int = 5
x: str = "five"
x
# => 'five'

No errors! While the type hints don’t match, Python won’t stop you. They’re hints, not guardrails.

Moving/Deleting Cells

Cells are independent entities, even if they can produce objects in the kernel that other cells can reference. You can reorder cells by clicking and dragging to the left of the code/text. You can also delete cells with DD when a cell is selected, or right-clicking on the cell.

Make sure to go through the related Notebook and complete the Check for Understanding before moving on to the next lesson.

1-2: Numbers

This is a quick one, but contains some important core concepts.

Data Types

All values in a programming language are of a certain data type. These types dictate how we interact with the information, and what properties and capabilities they possess in the language. We can use the type() function to determine the data type of any value or expression.

Types of Numbers

Let’s try that with the number 3.

type(3)
# => int

type(3) yields type int. It’s an integer, also known as a whole number. That’s distinct from the other kind of number in Python, a float. In a new cell, try:

type(3.3)

You get float. These are “floating-point” numbers, also known as number with a decimal point in it, and some degree of precision. But be careful: floats do not have perfect precision. Check this out.

This is a division operation. 10 divided by 3 is, of course, an irrational number. You might recall that the actual value should be 3.333…, with no end of 3s. Well, floating point numbers need to lop it off somewhere, and doing so has unexpected results at the tail end of its precision.

We often combine floats with the round() function, which take a value or expression, and a precision value. So round(10.0 / 3.0, 3) yields 3.333.

Basic Operations

Python has all the mathematical operations you might expect for numbers.

  • Addition: +
  • Subtraction: -
  • Multiplication: *
  • Division: /
  • Exponentiation: **

There is a square root function, but it’s in the math module, which we haven’t touched yet. However, if you remember your algebra class, you can achieve square roots by raising to ½, or 0.5.

100 ** 0.5
# => 10

There’s one less common operation you should know about: modulo.

Modulo (%)

The modulo operation yields the remainder of integer division. For example, 10 % 7 yields 3, because 10 divided by 7 in integer division (think long division) is 1, remainder 3.

That remainder is more handy than you might think! For example, it’s common to check for evenness or oddness of a value using modulo. n % 2 will always yield 0 for an even number and always 1 for odd.

Order of Operations

Good ol’ PEMDAS is respected in Python math expressions. But parentheses matter—both to the code and to human readers, so please don’t forget them!

1-3: Booleans

Our next data type is simple, but crucial. The bool type represents boolean values: either True or False.

Note

Boolean values are named for George Boole, a 19th-century English logician who pioneered logic through binary values.

True and False

The two variants of type bool are True and False. The capitals matter; those specific words are reserved for the bool variants. You can see this in action yourself by running the following expressions.

type(True)
# => bool
type(true)
# => NameError: name 'true' is not defined

Expressions and Comparisons

Most of the time, we don’t explicitly write out bool variants. Instead, we derive these values from expressions that evaluate to either True or False. An expression is any syntactically complete statement in Python that evaluates to a single value. 5 is a valid Python expression, but it does not evaluate to True or False. We need to test for truth or falsity.

Comparisons are a common way to perform such tests (but not the only way, as we’ll see later). Comparisons take two values, apply a test, and evaluate to True or False. You’re familiar with this even if you’ve never written a line of code. Here are the comparison operators in Python

  • ==: Equivalence (i.e. a is equivalent to b)
  • !=: Not equivalent
  • >: Greater than
  • <: Less than
  • >=: Greater than or equal to
  • <=: Less than or equal to

In practice this looks like:

7 > 5
# => True
2 <= 3
# => False
2 == 2.0
# => True
2 != 2
# => False

Negation

You see that != is the negation of the equivalence comparison. In some languages, ! is used to negate any boolean expression. In Python, we use the not keyword. Go ahead—put not in front of any of the comparisons you tried above. not flips the value.

Boolean Operators

Comparisons are useful, but we often need to evaluate more than one boolean expression at a time. Or we need to combine them for a different effect. That’s where the and and or operators come in.

These allow us to combine boolean expressions. For example, True or False evaluates to True. True and False will be False.

Just like in mathematical expressions, parentheses matter and affect order of operations. Try these two:

True and not False or True
# => True
True and not (False or True)
# => False

None

There is another data type closely related to bool known as NoneType, with a single value of None.

type(None)
# => NoneType

This is Python’s way of saying “Nothing’s here.” Other languages use null or nil to express this concept. I mention it here because None is “falsy,” meaning that it kinda evaluates as false. But only kinda.

None == False
# => False

None is not equivalent to False. But given an or statement between None and another value:

None or 10
# => 10

You’ll get the non-None value.

A cool way of seeing this is with a chained or expression. or expressions will evaluate to the first truthy value they encounter. So:

False or None or 6
# => 6

The first two values are ignored, but the int 6 is truthy, and therefore the result of the or expression.

Put simply, None will evaluate like False in comparisons.

1-4: Strings

We’ve technically already seen this data type, since we wrote "Hello, World!". Strings (str) are Python’s primary representation of text.

type("Python")
# => str

Strings can be in single quotes or double quotes, but for this course, we’ll standardize on double quotes.

Sequences

Strings are a type of data known as sequences in Python. We’ll see others shortly. Broadly, sequences are values made of smaller individual elements. In the case of strings, the elements are characters. Sequences are indexed, meaning the individual components can be accessed.

Indices

Sequences are indexed, like addresses on a street. The addresses begin at 0. So for the string Python, the indices go:

P Y T H O N
0 1 2 3 4 5

To access a specific index, you can use the square bracket notation.

msg: str = "Python"
msg[0]

# => 'P'

Indices can also go backward, which is the Pythonic way of accessing the last element of a sequence.

msg[-1]

# => 'n'

“Slicing” Strings

The square bracket notation actually has three potential parameters, separated by colons. The full syntax is:

seq[start:end:step]

Where seq is the sequence being indexed, start is the starting index, end is the ending index (non-inclusive), and step is the interval. We can use this syntax to “slice” the sequence. That’s a bit confusing until we put it all into practice. Let’s try an example with all three parameters used.

"0123456789"[1:8:2]

# => '1357'

We begin at index 1, the second character in the string. We continue up to index 8, but we move by steps of size 2, meaning we take every other character. The result is 1357. When the end parameter is omitted, it is assumed to be the end of the sequence. When step is omitted, it is assumed to be 1. You can also omit start and assume it to be 0.

step can also be negative.

"0123456789"[::-1]

That’s a common way to reverse sequences.

Concatenating Strings

Concatenation is the connecting or joining of items in a chain—or a string. We commonly want to build new strings out of pre-existing ones. Let’s say we have a variable name, and we want to create a greeting message. We can concatenate the name with a greeting to produce a final message. In Python, + is the string concatenation operator.

name: str = "Asmodeus"
greeting: str = "Welcome, "
message: str = greeting + name + "!"
print(message)

# => Welcome, Asmodeus!

So message is a string resulting from the concatenation of greeting, name, and an exclamation point.

Format Strings

Concatenation isn’t the only game in town for making new strings out of old strings. Format strings allow us to place variables directly inside of a “template string.”

We prepend the quotes of a format string with f to indicate it’s a format string.

Let’s revisit the greeting example with a format string.

name: str = "Asmodeus"
message: str = f"Welcome, {name}!"
print(message)

# => Welcome, Asmodeus!

Little cleaner, right? Fewer variables to keep track of. Whenever possible, I would use format strings when creating combination strings. They tend to be easier to maintain in the long run.

Oh, one more cool thing: the values inside of a format strong don’t have to be strings. They can even be expressions!

x: int = 5
y: float = 3.0
message: str = f"The sum of {x} and {y} is {x + y}"
print(message)

# => The sum of 5 and 3.0 is 8.0

1-5: Lists and Tuples

More sequences! As we mentioned before, strings are one example of a larger class of data types in Python called sequences. We’ll have to revisit one of the core characteristics of sequences in a later chapter, but we can introduce two more sequence types.

Lists

Lists are exactly what they sound like: a list of things. Just as strings are made of characters, lists are made of elements. Lists can contain any data type, and a mix of them.

Lists are surrounded by square brackets and each item in a list is separated by a comma.

stuff: list = [1, "two", 3.0, False, ["a", "nested", "list"]]

Yes, even other lists. Accessing items is identical to accessing characters in strings.

stuff[2]

# => 3.0

The same slicing syntax we saw for extracting parts of a sequence, or even reversing it, also works with lists.

So what’s new with lists? Other than containing more than just text, lists also possess special methods—built-in functions that allow us to interact with them.

Length

The built-in len() function measures the length of any sequence, including strings and lists.

len([1,2,3])
# => 3

Adding Stuff

The most common way to add to a list is with the .append() method. That adds an element to the end of a list.

# Did you think they would all be Trek?
x_men: list = ["Cyclops", "Jean Grey", "Professor X", "Wolverine", "Beast"]
x_men.append("Rogue")
x_men

# => ['Cyclops', 'Jean Grey', 'Professor X', 'Wolverine', 'Beast', 'Rogue'

There’s also .insert(), which will insert the value at a given index.

x_men.insert(3, "Gambit")
x_men
# => ['Cyclops', 'Jean Grey', 'Professor X', 'Gambit', 'Wolverine', 'Beast', 'Rogue'

Checking Membership

We commonly want to know whether a certain value is in a list. Python has the in operator for just such an occasion. We build in expressions like using other comparisons, with the element we’re checking on the left and the list on the right.

"Cyclops" in x_men
# => True

Note

The in operator works with all sequences—including strings!

Removing Stuff

The .pop() method removes the last item from a list and returns it. The return means the value is made available for assigning to a variable.

# Oh no! The Sentinels got the last X-Man!
captured: str = x_men.pop()
print(captured)
print(x_men)
# => 'Rogue'
# => ['Cyclops', 'Jean Grey', 'Professor X', 'Gambit', 'Wolverine', 'Beast']

.pop() can also be given a specific index to pop instead of the default last position.

wolverine = x_men.pop(4)
print(wolverine)
print(x_men)
# => 'Wolverine'
# => ['Cyclops', 'Jean Grey', 'Professor X', 'Gambit', 'Beast']

There’s also .remove(), which removes the first item (left to right) from the list that matches the provided value. That can get tricky when there are duplicate values. And remember, unlike .pop(), .remove() does not return the removed element.

x_men.remove("Cyclops")
x_men
# => ['Jean Grey', 'Professor X', 'Gambit', 'Beast']

Tuples

Another type of sequence closely related to lists is tuples. Instead of square brackets, tuples are surrounded by parentheses.

Tuples are immutable. Once you make them, that’s it. No .append(), no .pop(), no .insert(), no .remove(). They exist as created. Sometimes that’s exactly what you want though! And technically, tuples are less memory-intensive, since they don’t come with all that baggage. Here’s a simple tuple:

coordinates: tuple = (44.5898641,-104.7157436)
# Wonder where that is?

Tuples are still indexable, just not changeable.

You’ll frequently see tuples in the middle of workflows. Two built-in Python functions produce tuples: enumerate() and zip(). enumerate() takes a list and produces, well an enumerate object. That’s not super helpful yet, but we can make it a list for us to review.

list(enumerate(x_men))
# [(0, 'Jean Grey'), (1, 'Professor X'), (2, 'Gambit'), (3, 'Beast')]

Lots going on here. First, we see our first type conversion. You can use the name of a data type as a function to try to convert one data type to another. You’ll commonly see this with converting str to int or vice-versa. Here, we convert the enumerate object into a list so we can see what we made. The resulting list has the same number of elements as before, but each element is a tuple. Each tuple is a pair of the former list’s index and value.

zip() works similarly, but it takes two sequences and joins them together as tuples.

powers = ["telekinesis", "telepathy", "exploding cards", "is blue and furry"]
list(zip(x_men, powers))
# => [('Jean Grey', 'telekinesis'), ('Professor X', 'telepathy'), ('Gambit', 'exploding cards'), ('Beast', 'is blue and furry')]

A quirk of zip() is that the sequences don’t have to be the same size. zip() will stop producing items after the shorter sequence is exhausted.

list(zip(x_men, "abc"))
# => [('Jean Grey', 'a'), ('Professor X', 'b'), ('Gambit', 'c')]

1-6: Dictionaries

Let’s get this out of the way now: the Python data type we’re discussing is referred to as dict in code. But the full term is officially “dictionary.” Besides, “dict” is…not a great out-loud word. So in here, it’s a dictionary.

And in truth, the data type is a dictionary: the point is to look things up by a point of reference. Except instead of alphabetical order, we can cut right to the thing we’re looking for. It’s like a dictionary with every word easily bookmarked.

Imagine a list that was a million elements long. Now imagine we were searching for a specific element in that list. If the list has no inherent order (like sorted numbers), we have to just start at the beginning and inspect each element until we get lucky. In the worst case, that takes a million separate operations.

When we have data that we’ll want to look up specifically, dictionaries make more sense. They are a type of data structure commonly called key-value pairs. The “key”, is like a named address of the data, instead of a raw numerical index. The “value” is the data we want to access. Put another way, the key is the thing we know; the value is the thing we’re looking for.

Syntax

Dictionaries are surrounded by curly braces ({}). Keys and values are separated by colons, and pairs are separated by commas. It’s easiest to see in context. User data is a ready example for the value of dicionaries.

users: dict = {
  "admin": "password123",
  "jhackme": "summer2026",
  "apwnerson": "letmein!"
}

To access a given value, you reference the key. So to get Joe Hackme’s password, we’d put that username in square brackets after the dictionary name. Remember the key is a string, so we’ll need quotes.

users["jhackme"]
# => 'summer2026'

Adding/Removing Items

Dictionaries don’t really have .append() or .pop() methods the way lists do. Instead, to add an item, we simply reference a key with an equals sign. If the key already exists, the value will be reassigned.

Let’s add a user.

users["jphishman"] = "superman"
users["jphishman"]
# => 'superman'

Removing items is a bit trickier. If you want to blank out the key, you can assign it to None or and empty string (""). But that can get messy, since it introduces error possibilities in working with multiple data types and unexpected missing values. To delete a key-value pair, we have to use Python’s built-in del operator, which deletes the value.

del users["jphishman"]
users["jphishman"]

# => KeyError: 'jphishman'

You can use del with variables as well, but it’s rare that you’d want to.

Keys and Values

Sometimes you just want the keys of a dictionary. Or the values! Dictionaries have both .keys() and .values() methods that return collections of those, respectively

users.keys()
# => dict_keys(['admin', 'jhackme', 'apwnerson'])
users.values()
# => dict_values(['password123', 'summer2026', 'letmein!'])

In both cases, we don’t get back a list exactly. They’re special types called dict_keys and dict_values. But they are easily converted to lists with the same type conversion we’ve seen before.

list(users.keys())
# => ['admin', 'jhackme', 'apwnerson']

Checking Keys

Just like with lists, the in operator lets us check if a certain item is in a dictionary. When you use in with a dictionary, the check is against its keys, not its values.

"admin" in users
# => True

Nested Data

It’s common to see nested dictionaries, in which a dictionary’s values are themselves more dictionaries. That data structure mirrors what we often see in JSON (JavaScript Object Notation) data. Imagine a set of users where much more than just a password is keyed by username.

users: dict = {
    "admin": {
        "username": "admin",
        "password": "password123",
        "first_name": "Admin",
        "last_name": "Admin",
        "enabled": True
    },
    "jhackme": {
        "username": "jhackme",
        "password": "Summer2022",
        "first_name": "Joe",
        "last_name": "Hackme",
        "enabled": True
    },
    "jpwnerson": {
        "username": "jpwners",
        "password": "letmein!",
        "first_name": "Jane",
        "last_name": "Pwnerson",
        "enabled": False
    }
}

So to get the full record for the admin user, I’d use the "admin" key.

users["admin"]
# => {'username': 'admin',
# 'password': 'password123',
# 'first_name': 'Admin',
# 'last_name': 'Admin',
# 'enabled': True}

If I just wanted the password, I’d double up on the square brackets, accessing the "password" key of the dictionary retrieved at the top-level "admin" key.

users["admin"]["password"]
# => 'password123'

1-7: Conditionals

Python programs are instructions executed in the order written. At least so far, but that doesn’t really give us much flexibility. Sometimes we might want the ability to jump around in our program, to repeat instructions, or to only execute certain instructions under specific conditions. These abilities are known as control flow. Conditionals allow us to execute instructions only when a given test passes. Conditionals are the second of the 4 Fundamentals of Programming we’ve encountered.

Syntax

They look like this.

name: str = input("What's your name?")
if name == "Bob":
    print("Hey Bob! Good to see ya!")
else:
    print(f"Nice to meet you, {name}!")

Those few lines contain several important concepts and syntax patterns. First, we have the if keyword that starts the conditional statement. Then we have the condition, the test that evaluates to True or False, followed by a colon. The next line is indented. This is the code that will execute if the test passes (evaluates to True). It can be more than one line. After that comes the else block. Not every conditional statement has one of these; you can simply have an if that checks for a True state. This one has code that runs otherwise. And again, an indented line.

Note

Indentation is very important in Python. Other languages use curly braces ({}) to delimit blocks of code. Python uses indentation, which means it is both easier to write and easier to mess up.

We can handle more than one binary state. With elif, we can test for multiple conditions in order. Think of this as falling down a ladder. The first test to pass will be the code that executes, even if other conditions are also true.

num: int = int(input("What's the number?"))
if num % 15 == 0:
  print("FizzBuzz")
elif num % 5 == 0:
  print("Buzz")
elif num % 3 == 0:
  print("Fizz")
else:
  print(num)

Ah, good ol’ FizzBuzz. In this code block, providing a number divisible by 3 and 5 (also known as divisible by 15) will result in the first test passing. While one of the other conditions would also be true, that code does not execute because the first test captured the control flow.

Not every if needs an else, but if you use elif intermediate conditions, you must end with an else to catch all other possibilities.

Conditional Lab: Password Policy

If you’re here, you’re probably invested in defensive cybersecurity. It’s in the course title! For better or worse, some part of our lives as defenders concerns, yes, Compliance.

And Compliance loves a password policy, oh yes they do. We all know the words; let’s sing them together!

🎵 Your password must be… 🎵

🎵 Longer than eight characters 🎵

🎵 But not just alphabetical 🎵

🎵 Be sure to use a capital 🎵

🎵 And digits would be radical 🎵

🎵 And for that extra dash of strength 🎵

🎵 Use symbols to extend the length 🎵

🎵 If you follow these few simple rules 🎵

🎵 Your password won’t be guessed by fools 🎵

🎵 And that’s the password song!🎵

And now, let’s build a conditional tree that tests it!

# Get the password to test
password: str = input("What's the password?")

Length Check

The first check is for length. Python has a built-in len() function that will return the length of sequences, including strings. We can use that in our first test.

# Test for length
min_length = 8
if len(password) < min_length:
  print("[-] Length check failed!")

Mixed-Case

The second check is for mixed case (both upper and lower-case letters). For this, we can use the strings built-in lower() and upper() methods, which will lower-case or upper-case a given string, respectively. Our check can be whether the lower-cased version of the string matches the original or the uppercase version matches. Either way, it’s a fail.

# Use or to combine 2 checks
if password.lower() == password or password.upper() == password:
  print("[-] Mixed case check failed")

Digits

We need to check for the presence of digits. Now, while Python strings have isalpha() and isalnum() methods, we can’t just use those on our whole password and call it a day. For one thing, these tests will fail if we have special characters in there like we’re supposed to. Rather, we need to check to make sure that at least one character is a digit.

There’s an efficient way to do this, but it requires a trick we haven’t learned yet, so forgive the brief skip-ahead.

Python has a built-in any() function that returns True if any element of the sequence you give it evaluates to True. Watch:

# Any with booleans
any([False, False, False, True])

# => True

# Any with falsy values
any([0, 0, 0])

# => False

# Any with falsy and truthy values
any([0, 0, 1])

# => True

So what we need is a list of True/False values generated from the password string. We haven’t yet covered repeated actions, so forgive the brief leap, but we’ll use a list comprehension to apply isdecimal() to each character of password, and use any() on that list to determine if there are digits present.

Since we’re testing for their absence, we’ll negate the any() with not.

if not any([c.isdecimal() for c in password]):
  print("[-] Digit test failed")

Symbols

To seek symbols (naively), there’s a much easier method we can use. isalnum() returns True if every character in the string is alphanumeric. This should fail if symbols are present. It’s not perfect, but it’ll do for now.

if password.isalnum():
  print("[-] Symbol test failed")

Check for Understanding

Objectives

You have on objective: Combine these checks into a single conditional! Don’t forget we begin with if, have 0 or more elifs, and end with else.

1-8: Loops

Loops are yet another control flow structure, and number 3 in the 4 Fundamental Structures of Programming. Computers are all about automation, and automation is all about repetitive tasks. Programs can run as many times as we need, but often specific instructions within a program need to run multiple times as well. Enter loops.

In Python, loops come in two flavors. We have for loops when we have a pre-existing collection or sequence for which we want to repeat the same instructions. We also have while loops which will repeat the same instructions as long as a condition is True—or until a condition is False.

Syntax

We’ll start with the for loop. As I said, we can loop over collections, so let’s start there.

groceries: list[str] = ["milk", "eggs", "apples", "potatoes", "peanut butter"]
for g in groceries:
  print(g)
"""
=>
milk
eggs
apples
potatoes
peanut butter
"""

The for keyword begins the loop ceremony. We then create an iterator—a variable used to keep track of the current item we’re addressing in the loop. I like to think of this as counting on our fingers. Whatever element in the collection is next up in the sequence, that will be the value of the iterator. So for groceries, the first value will be milk. Then eggs, etc. After the iterator comes the in keyword, which tells Python to expect a collection. Then the collection. A colon and an indented line after, just like conditionals. Whatever is indented after the for line will be part of the loop.

A while loop works similarly, but does not take a collection. Instead, it takes a condition, and the iterator is tracked manually.

t_minus: int = 10
while t_minus > 0:
  print(t_minus)
  t_minus -= 1
print("Liftoff!")

Here, t_minus is the iterator. The while keyword expects an expression that evaluates to either True or False. In this case, a comparison of t_minus with 0. As long as t_minus is greater than 0, the loop will continue.

Note that inside the loop (and indented again), we use -= to manually decrement the iterator.

Warning

If you don’t handle the iterator, the condition will never change and the loop will never stop. This is a common mistake for new programmers.

range()

Sometimes you just need to loop through a series of numbers. In that instance, you use range(). This function produces a generator (more on those later) that will produce integers based on the arguments you pass to it. When given one argument, range() will provide integers from 0 up to that number. If you give it two arguments, it’ll be a start and stop. The start is inclusive while the stop is exclusive. Let’s demonstrate.

for i in range(10):
  print(i)
"""
=>
0
1
2
3
4
5
6
7
8
9
"""
for i in range(10, 20):
  print(i)
"""
=>
10
11
12
13
14
15
16
17
18
19
"""

range() can even take a third argument, which works like the interval component of slicing.

for i in range(10,100,10):
  print(i)
"""
=>
10
20
30
40
50
60
70
80
90
"""

A common pattern for using range() is when you need the index of a thing as well as the value. And we can pair this with len() to get the proper size for the range.

crew: list[str] = ["Janeway", "Chakotay", "Tuvok", "Paris", "Kim", "Torres", "Doctor"]for i in range(len(crew)):
  crewmember = crew[i]
  print(f"{i}: {crewmember}")
"""
=>
0: Janeway
1: Chakotay
2: Tuvok
3: Paris
4: Kim
5: Torres
6: Doctor
"""

You’ll see (and probably try to write) this pattern all the time. But remember, we have enumerate() to make this a bit easier. And when the elements of a collection are tuples, Python can handle a two-part iterator.

for i, c in enumerate(crew):
  print(f"{i}: {c}")

Isn’t that easier?

Building Collections

Frequently you’ll have to produce a string or a list based on a repetitive action, whether from another collection or from some iterative operation. These patterns are well-established, but there are more and less elegant ways to do it.

Iterative Building

In this pattern, we begin with an empty collection and concatenate/append everything we want.

# Squares of numbers
nums: range = range(1,100)
squares = []
for n in nums:
  squares.append(n * n)

# We don't need to see them all
squares[:10]

We start with an originating collection and a new empty one. Then we use a loop to populate the new list with the result of whatever operation we wish to perform on each element of the original. This is a tried-and-true pattern, and it is a perfectly valid way of building collections.

When building lists and dictionaries, Python gives us another pattern known as comprehensions.

List Comprehensions

Let’s produce the same list as above with a list comprehension.

squares: list[int] = [n*n for n in range(1,100)]
"""
=>
[1,
 4,
 9,
 16,
 25,
 36,
 49,
 64,
 81,
 100,
 ...
 9801]
"""

Same result, more efficient syntax. The comprehension is an expression wrapped in square brackets. Then, we have the operation first—whatever we want to do to each item in the collection. We need an iterator variable for that operation, so we do this a bit backwards. We use the iterator in the operation, then with the for keyword, we name the iterator. Then we state what collection we’re using with the in keyword.

It takes some practice, but once you get the hang of it, it’s an efficient to build lists. I use list comprehensions frequently in my code.

Dictionary Comprehensions

Turns out you can do the same thing with dictionaries, although the syntax can get a little trickier. Suppose we have a simple list of usernames, but we wanted a richer data structure with both usernames and passwords for each user? A dictionary comprehension can easily create a dictionary of keys from a given list.

users: list[str] = ["admin", "bob", "carol"]
users_dict: dict = {u: {"username": u, "password": "changeme"} for u in users}
users_dict

"""
=>
{'admin': {'username': 'admin', 'password': 'changeme'},
 'bob': {'username': 'bob', 'password': 'changeme'},
 'carol': {'username': 'carol', 'password': 'changeme'}}
"""

This comes in real handy when reshaping data during certain analysis tasks, especially when we start playing with tools like Pandas later. You’ll have a chance to practice the specifics, but for now keep in mind that both comprehension methods can quickly made new collections out of existing ones without the need for a full for loop structure.

1-9: Functions

We’ve arrived at the last of the 4 Fundamental Structures of Programming, and probably my favorite. Functions can be thought of as mini programs within the larger program that we can execute on demand. Anytime we have a task that must be repeated in a program, there’s an opportunity to make a function.

Basic Syntax

Let’s build a simple function to make our original “Hello, World!” code repeatable.

def greet():
  print("Hello, World!")

Short and sweet. def is the Python keyword that begins defining a function. Then comes the function name, followed by parentheses—in this case, empty parentheses, but that won’t always be so. Then a colon and an indented next line. Anything indented after the function declaration is part of the function.

If you created and executed the cell above, you received no output. That’s because declaring a function and calling or invoking a function are two separate operations, just like writing the program and running the program are separate. Another analogy I used to give my students was that declaring a function was like writing a spell down on a scroll, while calling the function was reading it aloud to cast it.

To call the function, we write the function name, followed by the parens and whatever the parens need to contain (hold onto that idea).

greet()

# => Hello World

Parameters and Arguments

What’s up with those parentheses, anyway? Just like how we can pass options to programs when we run them, we can provide information to functions to alter their output. When we define these in the function declaration, they are referred to as parameters. Think of them as variables we use inside the function whose value we can set when we call the function. Parameters make functions much more useful.

def greet(name: str):
  print(f"Hello, {name}!")

greet("Bob")

# => Hello, Bob!

Now our greeting has a name parameter of type str.

Note

As always, the type hints are optional. def greet(name): is also syntactically correct, but making parameter data types clear is one of the most important ways of documenting our code for other programmers.

We can name multiple parameters in a function. Too many gets cumbersome, but sometimes you want to provide more than one adjustable value.

def greet(name: str, punctuation: str):
  print(f"Hello, {name}{punctuation}")

greet("Alice", "!")
greet("Bob", ".")

"""
=>
Hello, Alice.
Hello, Bob.
"""

Now our greet() function is a lot more flexible. We can provide different arguments (specific values for function parameters) to alter the output.

Default Values and Named Parameters

In the greet() function, we have two parameters which are positional. We have to provide name first, then punctuation. But what if we didn’t feel like always providing punctuation, since most of the time it will be a period? We can set default values for parameters. This allows us to abbreviate the function call when we want to use the default.

def greet(name: str, punctuation: str = "."):
  print(f"Hello, {name}{punctuation}")

greet("Alice")
greet("Bob", punctuation="!!")

"""
=>
Hello, Alice.
Hello, Bob!!
"""

Naming parameters has an interesting side-effect. If, after our parameters without defaults, we only have parameters with default values, we get to change their order however we like. Let’s add a friendly boolean parameter to make the greeting more or less friendly.

def greet(name: str, punctuation: str = ".", friendly: bool = False):
  if friendly:
    print(f"Hello, {name}{punctuation} It's nice to see you{punctuation}")
  else:
    print(f"Hello, {name}{punctuation}")

greet("Alice")
greet("Bob", friendly=True)

"""
=>
Hello, Alice.
Hello, Bob. It's nice to see you.
"""

Notice I didn’t have to provide an argument for punctuation; I could just skip right over to friendly. This only works if we have exclusively parameters with defaults following any positionals.

Returning Values

Up until now, we’ve been using functions incorrectly. Our greet() function has been calling print() to generate output, but this short-circuits the intended usage of a function. Functions have a dedicated keyword, return, that sends a value back from the function. return ends function execution and sends the program back to the function call location. If we want to assign a value to a variable as a result of a function call, or really do anything else with the result of a function, the function must return a value.

Printing or any other operation inside the function that produces a change not represented in the return values is known as a side effect. We do not like side effects. Our functions’ operations should be clear and predictable based on the function signature (parameters and return type).

Let’s rewrite greet() to properly return a value.

def greet(name: str, punctuation: str = ".", friendly: bool = False) -> str:
  if friendly:
    return f"Hello, {name}{punctuation} It's nice to see you{punctuation}"
  else:
    return f"Hello, {name}{punctuation}"

greet("Alice")
# => 'Hello, Alice.'

Not much has changed—we’ve replaced print() with return. However, notice the subtle difference in the output. The value is now in quotes, because we’re returning a str. Also, did you notice the new type hint at the end of the function definition? We can use a -> type before the colon and after our parentheses to indicate the return type of the function. Like all type hints, these are optional, but I strongly encourage using them. Knowing what data types a function expects (parameters) and returns goes a long way to keeping track of your data in a Python program.

In Jupyter, returning values also has a somewhat annoying consequence for displaying values. In a cell, Jupyter will display the value of the last expression. If that’s a function call, no worries—you’ll get the returned value. But if you have two function calls at the end, you’ll only see the last one.

# In Jupyter:

greet("Alice")
greet("Bob")

# => 'Hello, Bob.'

To address this, you need to bring your own print() statements.

# In Jupyter:

print(greet("Alice"))
print(greet("Bob"))

"""
# =>
Hello, Alice.
Hello, Bob.
"""

Both values display. And notice the quotes are gone, because this is the result of printing output, not the display of the expression’s value. A subtle difference, but one that will become important as we handle and display multiple function results.

Function Composition

Some of you may remember this from Algebra. Since functions return values, a function call can be used directly as the argument to another function. This is known as function composition, and it’s an important design approach in many programs.

Let’s imagine a function that converts any string into SpongeBob case (wOrDs lOoK lIkE tHiS). It takes in a string and returns a string, and alters the case for each.

def spongebob(msg: str) -> str:
  result: str = ""
  for i, c in enumerate(msg):
    # lower case evens
    if i % 2 == 0:
      result += c.lower()
    else:
      result += c.upper()
  return result

What a fun demo function! We have a for loop in there, a conditional, and two string methods. This pattern of building a result over a loop and returning is one to get comfortable with. Also note we’re taking advantage of enumerate() to get the index and the value of each character in msg at the same time.

If we want to SpongeBobify a greeting, we can wrap the call to greet() in a call to spongebob().

spongebob(greet("Alice"))

# => 'hI, aLiCe.'

Check for Understanding

Time to write your own function. Remember the password policy from 1-7? Let’s functionalize it.

Objectives

  1. Write a function called valid_password() that takes a single str argument and validates it against our password policy. If the password is valid, the function returns True. If it fails, it returns False.
  2. Send the policy function (without parens) to testme(). Yes, you can pass functions as arguments!

1-10: Classes

Warning

This is going to be a long one. Get a drink or a snack before getting started.

We’ve just explored building functions and how to put them together. That can be a wonderfully elegant way to build programs, slowly changing information from one shape to another until we arrive at our destination. But that isn’t always the easiest way to think about problems. We’re humans, not data pipelines. We think about things. We can reason that way in Python too.

Python is an object-oriented language. “Object” has a specific meaning in here. It refers to data structures defined by a specification known as a class. Think of a class as a blueprint for that contain both properties and methods. Properties are characteristics, attributes that are common to a certain kind of structure, but whose values may be unique for a specific instance. Methods are functions attached to that object that take usually make use of the object’s specific property values.

Syntax

This makes more sense in practice. Let’s build a User class to contain information about users in a system.

class User:

  def __init__(self, username: str, password: str):
    self.username = username
    self.password = password

So much in so few lines. First, we use the class keyword to declare a new class, the name it. By convention, class names are capitalized.

Inside the class, we’ll see methods defined. Methods are like normal functions, except they take a special first argument, self, which refers to the specific object instance the function’s getting called on.

But how do you make instances in the first place? It’s all well and good to have a blueprint, but building from it is a different story. That’s what __init__() is all about. This function is known as a constructor. The double underscores (known as “dunders”) mark it as a magic method. These are methods Python will call under the hood in specific circumstances. This one makes object instances, and is invoked when we call the class name like a function.

The Constructor

The __init__() constructor method, like all methods attached to object instances, has a first parameter of self, referring to that specific object. Here we can use dot notation to set properties on the object. We are defining two properties (for now): username and password, both strings. We don’t need to do anything special with them, so they are passed right from the function to the property assignment. But that doesn’t have to be the case! Constructors can perform validation checks, sanitization, or other housekeeping before assigning properties.

Let’s make a new User instance called alice.

alice = User("alice", "password123")
alice
# => <__main__.User at 0x7f23f7eb3d90>

Properties

Asking for the value of alice doesn’t tell us much. Luckily, we can access the object’s properties.

Both properties and methods use dot notation for access from the object itself. We’ve already seen this at work with .append() on lists, and .lower() on strings. Both of those are methods common to all objects of the list and string classes, respectively. Let’s access our user’s properties.

print(f"{alice.username}: {alice.password}")
# => `alice: password123`

Methods

Wouldn’t it be nice if our user objects could do stuff too? Like, what if we wanted a way to easily reset the user’s password? If the User class had a defined reset_password() method, we could do it.

So let’s do it. We’ll rebuild our User class to add our new method, and reinstantiate alice.

class User:

  def __init__(self, username: str, password: str):
    self.username = username
    self.password = password

  def reset_password(self):
    new_pass: str = input("New password:")
    new_pass_confirm: str = input("Confirm new password:")
    if new_pass == new_pass_confirm:
      self.password = new_pass
    else:
      print("Passwords don't match")

# Recreate alice
alice = User("alice", "password123")

#reset password
alice.reset_password()

Note

Yes, I know the function we just wrote doesn’t return anything. Sometimes “getters” and “setters” will only make internal changes to the object.

This is a fairly contrived example, but now you see how we define and call methods on objects.

Magic Methods

You know, that format string we used earlier was pretty handy. It’d be nice if that was what we got when we asked Python to print() our Users. That’s doable! We simply need to define (and override) the __str__() magic method. That method, which is secretly already attached to our object (we’ll understand how shortly), determines what print() produces.

Let’s go back and once more redefine our User class a __str__() magic method.

class User:

  def __init__(self, username: str, password: str):
    self.username = username
    self.password = password

  def reset_password(self):
    new_pass: str = input("New password:")
    new_pass_confirm: str = input("Confirm new password:")
    if new_pass == new_pass_confirm:
      self.password = new_pass
    else:
      print("Passwords don't match")
      
  def __str__(self):
    return f"{self.username}: {self.password}"

Recreate the alice object and try to print it out.

alice = User("alice", "password123")
print(alice)
# => alice: password123

Much nicer.

Note

Try the __dict__() method. What does it do? Is that useful? And what about ints? Do they have magic methods?

Docstrings

Remember the help() methods and the ? trick in Jupyter? Try that on alice.

alice?
"""
=>
Type:        User
String form: alice: password123
Docstring:   <no docstring>
"""

Notice the “Docstring” section? Where do those come from?

Docstrings are multiline comments at the very top of modules (files), functions, and classes. By convention, they should either quickly describe what the function does, or do that and clearly define parameters.

Let’s once more redefine User, but with a docstring.

class User:
  """A User with a username and a password."""
  # ...

Reinitialize alice and use the question mark again:

alice?
"""
=>
Type:        User
String form: alice: password123
Docstring:   A user with username and password
"""

You’ll want docstrings for most classes, functions, and methods.

Class Methods and Static Methods

Most of the time, you’ll want the functionality related to a class to be object-specific. But sometimes, the class itself has some work to do. In that case, we can use class methods and static methods, alongside class variables

Class Variables (aka Class Attributes)

Let’s expand our User class to include an id property. We want this ID to increment every time we make a new user. That’s tricky, because how will the constructor know what the last id was? The objects can’t communicate that information between each other.

This is exactly what class variables are for. By attaching the information directly to the class itself, rather than a specific instance, the class structure can keep track of the information throughout the run of the program. Let’s rebuild the User class one more time with a class variable and a new property for the instances.

class User:

  # Our new class variable
  last_id: int = 0

  def __init__(self, username: str, password: str):
    self.username = username
    self.password = password

    # Increment last_id and use it for our id
    User.last_id += 1
    self.id = User.last_id

    # A cleaner output for our users
    def __str__(self):
      return f"""ID: {self.id}
Username: {self.username}
Password: {self.password}"""

    # Rest as before...

So we’ve added last_id to the User class itself. When __init__() runs, it increments that value and uses it for the id of the newly-constructed User object. The next time the constructor runs, User.last_id will have that new value ready and waiting.

While we were making changes, I went ahead and added the id value to the string representation of the object. Notice that I used a multiline string, making a nice clean output for out information.

Try it out. Make two users and see how they look.

alice = User("alice", "password123")
bob = User("bob", "letmein")

print(alice)
print(bob)

Sure enough, the ids increment.

Static Methods and Class Methods

Methods we’ve seen so far attach directly to the objects we make. But there are situations in which we need helper functions that don’t have to exist on every object, but still make sense within the class container.

If such a helper function needs access to data stored in the class (aka a class variable), we will create a class method. If it’s just a utility function that doesn’t access class variables, we will make a static method. What’s the difference?

Two things: a parameter and a decorator. Let’s make both and you’ll see.

First up, a class method to make a blank reference user without updating the last_id. Modify User by adding the below function, including the @classmethod line.

class User:

  #...
  @classmethod
  def reference_user(cls):
    new_user = User("notauser", "Password123")
    # Re-decrement the last_id after the constructor incremented it
    cls.last_id -= 1
    return new_user
  #...

New syntax! The @classmethod decorator is a sort of macro for the function. It allows us to define a function inside the class that doesn’t get interpreted as an instance method, which is what Python expects all functions in a class to be otherwise.

The function itself includes a cls parameter, which references the containing class. That’s how we can access cls.last_id.

Note

Why not just reference User.last_id? In this case you could, but as we’ll see when we get to inheritance, you might not want to hard-code a class name that way.

Now let’s add a simple password validator. We won’t make the rules too complicated; the important point here is that we have a utility function that makes sense underneath the User namespace, but doesn’t actually need access to any instance- or class-level information. It’s just…in there. For that we’ll use a static method. Add the following to our User class.

class User:

  #...
  @staticmethod
  def validate_password(password: str) -> bool:
    """Returns true if password is valid per policy:

    1. >= 10 characters
    2. Contains capitals
    3. No spaces
    """
    return len(password) >= 10 and \
      password.lower() != password and \
      " " not in password
  #...

Now our User class has a handy validate_password() function we can access at any time.

User.validate_password("password")
# => False
User.validate_password("Password123")
# => True

Inheritance

We sometimes encounter the need to distinguish between different variations on a class. They’ll all have some things in common, but details will differ. Users can reflect this variation, since a system may have many kinds of users—they’ll share some characteristics, but also differ in ability and sometimes even stored data.

We can represent these different user types with User of User. In fact, we can build our classes such that User mostly functions as a base class that can be extended or modified for more specific user types. Let’s first amend User to include a role property, and adjust the __str__() method accordingly.

class User:
  # ...
  def __init__(self, username: str, password: str):
    self.username = username
    self.password = password
    self.role = "user"
  # ...
    def __str__(self):
      return f"""ID: {self.id}
Username: {self.username}
Password: {self.password}
Role: {self.role}"""
  # ...    

Now we can create an Administrator user type that inherits from User. Inheritance is indicated with parenthesis after the class name in the declaration, and the parent class or super class within.

class Administrator(User):
  """Administrator User"""

  # Stricter controls for admins
  @staticmethod
  def validate_password(password: str) -> bool:
    """Returns true if password is valid per policy:

    1. >= 20 characters
    2. Contains capitals
    3. Contains special characters
    3. No spaces
    """
    return len(password) >= 20 and \
      password.lower() != password and \
      any([c in password for c in  ["!\"#$%&\'()*+,-./:;<=>?@[\\]^_`{|}~"]]) and \
      " " not in password


  def __init__(self, username, password):
    super().__init__(username, password)
    self.role = "administrator"

Two changes here, other than the declaration using the User class as a parent in the parentheses. First, we override the validate_password method for the Administrator class, because this class has a stricter password policy. Second, the constructor first calls super(), which returns the parent class, and invokes its constructor with the arguments passed to Administrator.__init__(). Then we modify the role separately for admins.

Let’s make a new admin user just to prove it works.

admin = Administrator("admin", "ThisIsASuperSecretPassword123!")
print(admin)
"""
=>
ID: 1
Username: admin
Password: ThisIsASuperSecretPassword123!
Role: administrator  
"""

This is how we can create variations of complicated objects that share significant amounts of structure.

There’s some art to thinking in objects when writing code, and not everyone agrees on the best ways to do it. Some people avoid it altogether and stick to writing lots of composed functions. But when a problem naturally fits into object-shaped patterns, you should have the tools to express those ideas clearly, and to make distinctions between super and sub-classes of objects.

Real World Application: The Indicator Class

We’re going to practice with one more class design, this time pulled directly from the cybersecurity world: indicators of compromise (IoCs).

Atomic indicators of compromise like URLs, domain names, and IP addresses are not always high-fidelity, but will be routinely available and worth using in many situations. They’re the easiest to block, even if they’re also the easiest for the attackers to change up. Sharing these can be a little tricky because we have safety standards concerning “defanging” these indicators. For example, a URL will have its scheme modified and the last dot of the domain bracketed to prevent accidental clicking. So https://taggartinstitute.org becomes hxxps://taggartinstitute[.]org.

To represent indicators in Python, we’ll make a base Indicator class, then 3 subclasses for IPv4Indicator, URLIndicator, and DomainIndicator. We’ll also create a defang() method that may be handy for safely passing indicators around to teammates and documentation.

To start, let’s build the base class.

class Indicator:
  """
  Atomic Indicators of Compromise.
  
  Parameters
  ----------
  
  value: str
      Value of indicator
  """
  
  def __init__(self, value):
    self.value: str = value
      
  def defang(self) -> str:
    """
    Defangs the indicator
    
    Implemented in subclasses
    """
    pass

Not a lot going on, but it’s a start. Note that we have pass for the defang() method. This is Python for “Haven’t done this yet,” which is fine for our base class. Now let’s build one of our subclasses.

To properly inherit, not only will we need to add something to our class line, but we’ll also need to take advantage of the built-in super() function, which returns the parent class. From there, we can access its __init__() constructor and pass in our own constructor’s arguments to inherit properties without reinventing them.

We can also properly implement defang() on our child class in a way that’s appropriate for IPv4s.

Defanging IPv4

Let’s think about this for a moment. It would be simple to just replace every . with [.] in the value, but that’s not actually what the convention is! To refang every dot would be a huge pain. Instead, we just want to bracket the last dot. How to do that?

Well, strings have two methods that can help: .find() and .index(), which both return the earliest index of a given substring. Of course, we want the latest index, not the earliest. The solution? Reverse the string, find the first dot, replace it, and reverse it again.

It sounds more complicated than it is.

# Notice the parens? Parens for Parents!
class IPv4Indicator(Indicator):
    
  def __init__(self, value):
    # Instantiate parent properties/methods
    super().__init__(value)
      
  # Overwrite the `defang()` method for our purposes
  def defang(self) -> str:
    """
    Defangs the indicator
    
    Brackets the dots
    """
    
    # Reverse the value
    rev = self.value[::-1]
    # Find the first dot — Errors out if none
    dot_idx = rev.index(".")
    # Reassemble the string with slicing. Take until the dot, add the defanged dot, then add the rest
    # of the address. Note we have to flip the brackets for the reversing, and the second slice
    # starts *after* the existing dot
    defanged_rev = rev[:dot_idx] + "].[" + rev[dot_idx+1:] 
    # Reverse the reverse for the final return
    return defanged_rev[::-1]
i = IPv4Indicator("1.2.3.4")
i.defang()

Let’s do one more together, then it’s up to you. We’ll do DomainIndicator.

class DomainIndicator(Indicator):
    
  def __init__(self, value):
    # Instantiate parent properties/methods
    super().__init__(value)
      
  # Overwrite the `defang()` method for our purposes
  def defang(self) -> str:
    """
    Defangs the indicator
  
    Brackets the dots
    """

    # Yeah, domain defanging and IPv4 defanging are the same, basically.
    # Reverse the value
    rev = self.value[::-1]
    # Find the first dot — Errors out if none
    dot_idx = rev.index(".")
    # Reassemble the string with slicing. Take until the dot, add the defanged dot, then add the rest
    # of the address. Note we have to flip the brackets for the reversing, and the second slice
    # starts *after* the existing dot
    defanged_rev = rev[:dot_idx] + "].[" + rev[dot_idx+1:] 
    # Reverse the reverse for the final return
    return defanged_rev[::-1]

Can you complete a URLIndicator yourself?

1-11: Modules

So far, we have stuck to Python code that we’ve written ourselves inside these Notebooks—or whatever’s built into Python. But there’s a wide universe of Python code out there to explore and take advantage of. We may even want to package some of our code the same way. That’s why we need to understand Python modules.

Simply put, modules are packages of Python code that we can import into our project or notebook.

We’ll begin with an out-of-the-box Python module, like random.

Syntax

Python has plenty of additional modules, which you can review in the Library Reference. We’ll play with random for the moment.

To import an entire module, we can simply import the module name.

import random
random?

The random module is useful for, y’know, random stuff. Like random integers/floats, or even a random choice from a sequence.

When the entire module is imported, its constants and functions reside under the random namespace. That means to access the choice() function, we would refer to random.choice().

Note

Even though this uses dot notation like classes, random is not a class. It’s just that Python uses the same syntax to access the namespace of modules and classes.

# A random, namespaced choice
random.choice(["Klingons", "Romulans", "Borg", "Oh my!"])

# => ???

But importing random brings the whole module along. We may want to be more specific with our imports. To do that, we can use the from...import syntax. Instead of importing everything, we just import what we want. When constants and functions are imported this way, they are not namespaced—instead, they are directly available.

# Instead of importing everything, just get what we need.
from random import choice

# Look Ma! No namespace!
choice(["Klingons", "Romulans", "Borg", "Oh my!"])

Installing Modules

As rad as Python is out of the box, you’ll likely want to install somebody else’s code eventually. It’s not hard! In fact, you can even do it from directly within a Notebook if you want.

This is kind of an aside, but a cool trick that Jupyter can do is executing shell commands from within a notebook. All it takes is prepending the command with a ! (bang). Watch:

# List the contents of our repo
! ls ../

We can even save the result to a variable!

stuff: list = ! ls ../
stuff

Amazing, right? More on that later. But for our purposes, that means we can install packages directly from notebooks. If you’re not in a virtual environment, you might just use the pip3 package manager directly. But we are in a virtual environment. What’s more, we’re using uv. We can either run uv add or uv pip install to add the package to our venv.

! uv add requests

In our case it was already there, but the principle is sound. Once the package is installed, we can import it.

# A very naive web request
import requests

r = requests.get("https://taggart-tech.com")
r.text

# => A lot of HTML

Creating Modules

Not all our Python code has to live in the Notebook. In fact, it’s probably a good idea that most utility functions, classes, etc. live elsewhere. A good rule of thumb is: if the code should be modified by users on each run of the notebook, leave it in. If it’s small and helps the reader/runner understand the Notebook, also leave it in. If the code never changes and just needs to exist for the Notebook to work properly, get it out.

Luckily, basic module creation is as simple as making a new .py file and chucking your code in there. You can them import with import filename without the extension. Just try not to use the same filename as a real module you’ll also want to import, or a function you’re already using. For example, do not under any circumstances make a print.py file.

import print

# Sure you have a stuff function in there, but AT WHAT COST?!
print.stuff()

print("This won't work now")

You’ll have a bad time.

__init__.py

Hey look! Dunders again!

If you want to make a folder full of Python files but only import with a single statement, it can be helpful to use the Package structure. There’s a lot to this, but broadly, __init__.py inside a folder makes Python treat the folder as a module. Any .py files inside that folder can be imported via dot notation.

Imagine I have a folder call mymod and a file inside called foo.py. And inside that file is a function called bar(). If the folder had a blank __init__.py, I could import that function with:

import mymod.foo

mymod.foo.bar()

But what if I wanted to make it easier to access my submodules?

Alternatively, I could bring the submodules into the top namespace with from...import:

from mymod import foo

foo.bar()

Lastly, if wanted to defined what importing everything from my module really meant, I would add a definition of the magic __all__ list to my __init__.py.

__all__ = ["foo"]

This allows me to import everything defined in that list with *.

from mymod import *

foo.bar()

1-12: Errors

Before we move on to the meaty stuff, I want to spent a moment on one of the most important skills you’ll develop as a programmer: reading and handing errors.

Errors are going to happen in your code; that’s normal and expected—even if at the time they’ll make you want to pull your hair out, question your intelligence and purpose in life, right up until you find the stupid bug. Programming is a humbling pursuit, because the computer will not do what you intended, only what you wrote. We’ll leave aside the inference engines that devour meaning for this exercise, choosing instead to focus on the joy of making something ourselves—including mistakes.

Let’s make a mistake together.

# An intentional mistake
x: int = "10" #Whoops
x / 2

# => TypeError: unsupported operand type(s) for /: 'str' and 'int'

Something went wrong, hooray! Don’t panic. The first thing to remember when errors occur is that you can fix them. And the error message is your best chance to do so.

A Python error, also called an Exception, can come in many different flavors. Read the docs on Exceptions to see all the different kinds that are built-in. Third-party modules will have their own as well. In this case, we have a TypeError, and a message telling us that we used an “unsupported operand” with /. Sure enough, you can’t divide a string.

You can also see that the traceback points you right to where Python thinks the error is. This isn’t always perfect, but it gets you close (sometimes the issue is a missing comma/indentation above or below).

Handling Errors

It’s all well and good to understand error messages, but we don’t want to have to debug a crashing program in production. Good programming is defensive, meaning it anticipates and handles potential errors or bad input gracefully. Python gives us a valuable structure for doing so: a kind of control flow known as try/except:

The try/except block attempts an operation, and then provides “escape hatches” in the event of failures. Multiple types of exceptions can be handled in different branches, and we can also provide alternative instructions if none of those match. We can even provide instructions that run regardless of failure state.

Let’s build a common error pattern to test this out: accessing an out-of-range index in a collection.

# The empty list
stuff = []
try:
  thing = stuff[1] #That won't work
except:
  print("That index doesn't exist!")
# => That index doesn't exist!

Seemingly simple, but what’s important here is that the exception was handled. The program did not crash.

Let’s get more specific with our exception handling—one for a bad index, another for a bad list name.

stuff = []
try:
  thing = stuff[0] #Change the list name to something else to see the difference
except IndexError:
  print("That index doesn't exist!")
except NameError:
  print("That item doesn't exist")

We’re making some guesses about errors to expect, but we may not be able to anticipate them all. We can also name the Exception and refer to it in our code.

stuff = []
try:
  thing = stuff.keys() # Accessing a non-existent attribute
except IndexError:
  print("That index doesn't exist!")
except NameError:
  print("That item doesn't exist")
except Exception as e:
  print(e)

# => 'list' object has no attribute 'keys'

We hadn’t named that AttributeError, but we could use it in our catch-all except, which triggered because the more specific except clauses above did not.

We can also add code that runs if the attempt is successful in an else block, which must come after all excepts. This is a useful pattern that allows us to isolate the error-prone code in the try. You don’t want too much in the try, or you may catch errors from another step that needs separate handling.

stuff = ["vanadium", "iridium"]
try:
  thing = stuff[0] # Accessing a non-existent attribute
except IndexError:
  print("That index doesn't exist!")
except NameError:
  print("That item doesn't exist")
except Exception as e:
  print(e)
else:
  print(f"Picking up some {thing}")


# => Picking up some vanadium

Finally, there’s a finally! The finally clause allows us to execute code regardless of exception state. This can be useful for logging and other operations that should occur in all circumstances.

Mess with different parts of the code below to see how the output differs.

stuff = ["vanadium", "iridium"]
try:
  thing = stuff[0]
except IndexError:
  print("That index doesn't exist!")
  status = "index_error"
except NameError:
  print("That item doesn't exist")
  status = "item_error"
except Exception as e:
  print(e)
  status = "error"
else:
  print(f"Picking up some {thing}")
  status = "ok"
finally:
  print(status)

There you have it, a quick and dirty introduction to errors in Python. Learn to love try/except—it really will save you lots of time and frustration. One area that’s particularly true is in long-running procedures in Notebooks that handle big datasets. You do not want to be halfway through thousands of items and have something error out, requiring you to start again. Anticipate the errors, handle them in code, and run the process with confidence.

2-1: Markdown

Welcome to Part 2! We begin by discussing the most important part of any Jupyter Notebook: the writing.

Although we’ve seen Markdown in every notebook in Part 1, we haven’t really dug into the how and why of using these cells in a Notebook. There’s an art to it. You want to create enough context to explain what’s going on in the Notebook, while still letting the output speak for itself as much as possible.

As a reminder, you can get a fast, but thorough, introduction to Markdown syntax here

Pro tip: to make the most of this Notebook, you should be double-clicking on each cell to see how it was made in Markdown!

Headings

Different levels of headings (h1-h6) in Markdown make more of a difference than you might think. In addition to visually organizing your Notebook for reading, it allows the generation of a table of contents. Look over on the left for this icon:

See that? Yeah, Jupyter automatically creates a table of contents based on headings.

There’s also another trick Jupyter can do with headings: create slides. We’ll demo how that works in a later section, but it is incredibly handy for presenting data to others.

TL;DR: use headings. They may save your life!

Images

I just needed another spaceship

Images can be added to Markdown cells a variety of ways: image links from the internet, local file inclusions, and even pasting from the clipboard! That last one tends to be the source of a lot of the images I use in a Notebook, like the ToC icon in the previous cell.

Although Galactica sure is pretty, don’t overuse images in your Markdown cells. Use them where appropriate, like for flowcharts to explain a process, or instructional screenshots to help the Notebook runner understand what’s taking place.

My rule of thumb is to try to minimize the amount of scrolling one has to do in a Notebook, and images take up a lot of space. So if you use them, make them worthwhile.

Although you want your readers/runners to stay on the Notebook as much as possible, often it’s necessary to link to external datasources or references. You have a few choices in how to do so.

Inline

For informal references or links to documentation to help the Notebook runner, inline links like this are just fine. Try to structure your sentence such that the words you make hyperlinks clearly indicate what the link will be. I didn’t do that in the first sentence of this paragraph. But if I wanted to talk specifically about the functools module and all the cool capabilities therein, the link makes a lot more sense.

In-Cell Footnotes

Unfortunately, although some flavors of Markdown allow for footnotes with a special syntax, Jupyter does not support this readily. It does, however, support <sup> tags for superscript characters.1 You can even use the automatically generated link for a heading to send the reader to your Notes section when clicking on it.

Ending Notes

Alternatively, you can add an entire Notes cell with referenced by superscript annotations. This may be more appropriate for academic writing, allowing for a proper Works Cited/Bibliography for your work. In my opinion, this format should only be used for formal citations, as otherwise it completely disrupts the reading/executing flow of the Notebook. As much as possible, keep your readers’ eyes on the cell in question.

Notes

1: See?

Admonitions

Warning

This is an admonition

Note

So is this.

Admonitions are not a standard Markdown feature, but one you’ll find in Github-Flavored Markdown, a common dialect. We’ve installed the jupyterlab-myst package to take advantage of it. They’re handy for calling attention to important points in your document. But don’t overuse them. or they lose effectiveness.

Small Code, Big Text

That’s the rule to write by when creating Jupyter Notebooks. Be detailed (but not overly verbose) in Markdown, and succinct in your code. Take advantage of modules to abstract away complex processes that the reader/programmer doesn’t need to see.

In an academic Jupyter Notebook, our intention is to demonstrate method clearly and create reproducible results. But in a professional Notebook, we care much more about the latter. Explain what’s happening clearly, then get it done efficiently.

That does it for this gentle introduction to Part 2. Keep these principles in mind as we continue through the course. In the next lesson, we’ll go over one more Jupyter trick to make Notebooks more user friendly: Widgets

2-2: Widgets

2-2: Widgets

Notebooks are a kind of application. Granted, they don’t have a traditional user interface, but in every other sense, they are an app: code repeatedly used by a person to perform a task, structured an organized for usability.

It will frequently be the case that users will have to provide variable inputs like uploaded files, text, etc. for the Notebook to run appropriately. While we can use raw variable assignments or input() calls, we have a better way: ipywidgets.

Let’s get them imported.

# Import widgets and give them a nice name
import ipywidgets as widgets

Basic Usage

Widgets allow us to provide easy-to-use UI components in a Notebook for changing common parameters. But from a code perspective, they get a little funky. Lemme show you. We’ll start by making a slider that can set an integer value.

widgets.IntSlider(min=0, max=100)

Cool, right? But how do you actually get that value and use it? That’s where it gets a little trickier. In order to access the value within the object created by Widget methods, we need to assign the widget to a variable. But when we do, the widget no longer actually displays. For that reason, we need to also import another function from the IPython.display module: display().

Note

Fun Fact: IPython, or Interactive Python, was Jupyter’s precursor!

# Import display
from IPython.display import display

Now that we have the display() function, we can both show our widget and grab the value.

w = widgets.IntSlider(min=0, max=100)
display(w)
# Print the value of the widget
print(w.value)

By the way, have you clicked on that sidebar next to the cell yet? With it or the View->Collapse options, you can hide code or outputs for a given cell. For widgets, that means you can just show the widget without having to show the code that makes it!

Other Widgets

Let’s explore some of the other widgets and how they work. This will by no means be a comprehensive review of every widget. I encourage you to review the Widget List from the Jupyter Widgets documentation, as well as the rest of those docs. But for now, enjoy these samples!

Progress Bar (IntProgress and FloatProgress)

The Progress Bar is a fun way to allow a user to track the status of a long-running process. This also provide an opportunity to introduce the concept of named parameters in functions. We’ve already seen min and max for the IntSlider. Functions can take position parameters, identified by their order in the function call, or named parameters, which can be identified by name and have default values. For IntProgress and FloatProgress, we have a number of options to set. Check them out.

# Look at all the options we can set!
progress_bar = widgets.IntProgress(
    value=0,
    min=0,
    max=100,
    description='Loading:',
    bar_style='', # 'success', 'info', 'warning', 'danger' or ''
    style={'bar_color': 'magenta'},
    orientation='horizontal'
)
display(progress_bar)

We can use the time.sleep() function to simulate a process taking time, and updating the bar’s status to reflect that.

# Set initial value for each run of this cell
progress_bar.value = 0
progress_bar.bar_style = ""
progress_bar.description = "Loading:"

# Import our sleep() function
# sleep(n) pauses execution for n seconds
from time import sleep

# Add 10 every second until we reach 100
while progress_bar.value < 100:
    progress_bar.value += 10
    sleep(1)

# Change the style and text for the bar upon completion
progress_bar.description = "Complete!"

When you have multiple options to choose from, you in fact have multiple widget options to choose from! One is the Dropdown widget. Works as you might expect:

# Create the options list (can do this inline, but I like this more)
ships = ["Enterprise", "Hood", "Reliant", "Excelsior", "Grissom"]

# Create the dropdown
dropdown = widgets.Dropdown(
    options=ships,
    value='Enterprise',
    description='Select a Ship:',
    disabled=False,
    layout = {"width": "max-content"}, # For long item names,
    style={"description_width": "max-content"} # For long description labels
)
display(dropdown)
print(f"You selected {dropdown.value}.")

Buttons

Running cells is great and all, but this is the 21st century! The age of the atom! We’ve walked on the moon! Cars drive themselves very poorly!

Anyway, since we’re in the pushbutton age, why not…push some buttons to trigger events? With the Button widget, you can connect a Button to a function you’ve defined. That means you can run code with a more controlled click than running cells raw.

Now, to show output from click events, we need to use another widget as well: The Output widget. Additionally, you’ll notice a new syntax here: with. This is syntactic sugar that simplifies common open/close procedures. In this case, we use it to use the Output widget as the destination for a print() invocation.

Note that we have to display both the Button and the Output widget for this to work! UI programming is fun, innit?

Play with this one!

# Define the Output
output = widgets.Output()

# Define the button click handler.
# Note the btn argument, which the button will provide to the on_click function
def button_event(btn):
    with output:
        print("I came from a button!")

# Define the Button
button = widgets.Button(
    description="Click me",
    disabled=False,
    button_style='', # 'success', 'info', 'warning', 'danger' or ''
    tooltip="Click me",
    icon="atom", # (FontAwesome names without the `fa-` prefix)
)

# We have to display both the Button and the Output
display(button, output)


# Set the on_click _AFTER_ displaying
button.on_click(button_event)

Check for Understanding

The nature of this check means we can’t do an automatic test, so please do your best to complete this challenge!

Objectives

  1. Using Widgets either described here or from the Jupyter Widgets docs, create a small app that uses our Indicator module from Part 1.
  2. Allow the user to input text, then select whether it is a Domain, IPv4, or URL
  3. Have a “Defang” button that returns the defanged value.

You will almost certainly have to consult the external docs to make this happen. Take your time and use this opportunity to practice with everything we’ve covered so far.

2-3: File Upload

The File Upload widget is complex enough (and important enough!) that I felt it merited its own lesson.

Uploading files to a Notebook allows users to dynamically change what material the Notebook works on. This can be more ergonomic than alternatives like creating a folder where users dump input files. Then again, you may not want to use the File Upload widget if the expected files are very large. Leave those outside the kernel’s memory!

Now, before we explore the file upload widget, we have to cover a new kind of data we’ve not discussed before: bytestrings.

Bytestrings

Strings, as we’ve explored, are sequences. Sequences of what? What is it that strings are really made of? Tinier strings?

Kinda, but this is computing after all, so at some point text is in fact a representation of a number somewhere. Check this out:

[ord(c) for c in "I'm made of numbers!"]

ord() show us the numerical representation of each character in our string. This is possible because each character is in fact stored as an 8-bit integer.

Note

I know there’s more to it than that, but for now the simplification is useful._

8 bits make a byte, and so each character is a single byte of data.

This table from asciitable.com shows the entire table of ASCII characters.

The mathematically inclined among you may notice that there are only 128 values here, but that an 8-bit integer goes up to 255.

That’s true. The extended ASCII table shows us the rest!

We’ll leave aside unicode for now.

Point is, each character is in fact a number in disguise! But you might also notice that some of those characters (0-32) are not “printable” in the normal sense. But we need some way of representing them! That’s what bytestrings are all about. They give us a way of showing all the characters in a sequence of bytes, not just the ones that we can print normally.

Just as ord() gives us the numerical representation of a string, chr() converts a number to a character. We can use this to test what happens for different values.

Let’s start by getting the numerical value for A:

ord("A")

And now, let’s convert it back:

chr(65)

Simple, right? Now let’s try one of those unprintable characters like, 0.

chr(0)

Whoah! What just happened? What is that? It’s a string, but chr() added some stuff to it! That’s a string representation of a bytestring! But how can we make a bytestring directly?

Just like int() and str() there is also a bytes() that will create a bytestring, and also show us the easy way. bytes() is cool because it takes just about any kind of input.

bytes?

Let’s try it first with the iterable (list) of integers.

# A single byte
bytes([65])
# Two bytes
bytes([65, 0])

A lot just happened. First, we created a new bytestring out of a single integer, which became b'A'. That b is the secret easy way to create bytestrings! We can use that to make our own without the bytes() constructor.

# Create a bytestring quickly
bytestring = b"I'm made of bytes!"
print(bytestring)
type(bytestring)

When we gave bytes() the 0 as well as the 65, we got back b'A\x00. The \x is Python’s way of saying, “Hey this isn’t printable, so recognize this as the numerical representation of that character.” Why \x? Because the number is in hexadecimal.

We’re going to assume at this point that you’re familiar with hexademical values. But if you need a refresher, check this out.

Encode/Decode

You might have noticed that if we give bytes() a str, we also have to give it an “encoding,” whatever that is. Basically, it’s a translation table that informs Python how to represent the bytes as characters. Normally, we’ll use utf-8 for Unicode Transformation Format, 8 Bit (Extended ASCII). Let’s try it:

# Create a bytestring from a str by providing an encoding
bytes("ABC", "utf-8")

Moving between str and bytes is pretty common in Python, and in fact why we took this little detour.

The important thing to remember is that encoding produces a bytestring and decoding produces a str.

bytes objects have a decode() method, and strs have an encode() that defaults to utf-8.

# Move from bytes to str via decode()
b"\x41\x42\x43".decode()
# Same thing, but with regular character representation
b"ABC".decode()
# And back to bytes
"ABC".encode()

Actually Using the File Upload Widget

You might be wondering why we spent SO MUCH TIME talking about bytes in this lesson about the File Upload widget. In addition to it being a good time to discuss bytes…

The reason we went through all of this is because The File Upload Widget Stores file contents as bytes!

So with an understanding of bytes, we can finally use this thing.

# Import what we need
from IPython.display import display
import ipywidgets as widgets

You’ll notice that there’s a sample indicators.txt in the Notebook folder for you to upload. Toss it in there!

# Create the File Upload widget
upload = widgets.FileUpload()
display(upload)

⬆️ That’s what all this fuss was about! Can you believe it?!

The File Upload widget has 2 optional parameters: accept, which takes a str of comma separated MIME types or file extensions like .txt to accept; and multiple, which defaults to False, but when made True allows the widget to accept multiple upload selections.

With that file uploaded, let’s check out the widget’s value:

# Get the upload value
upload.value

What are we looking at here? The data is a nested dictionary, with a single top-level key: the name of the uploaded file. If we had uploaded multiple files, each would have its own key.

Inside that key are two more: metadata and content. The metadata key contains yet another dictionary of information about the file. At long last, the content key contains the file’s contents. In what format?

Well well well, look at that: a bytestring.

But now, we know how to get the data as a string!

# Get the file contents as a string
indicators: str = bytes(upload.value[0]["content"]).decode()
indicators

Cool, so we have a string. But when we pull in a multiline file as a (byte)string, the line breaks become \n characters. That might be okay, but commonly we’ll want each line as its own thing, meaning converting the single string into a list. Luckily, strs have a built-in method, split(), which can turn a single string into a list by separating on a specific pattern.

# Get the file contents as a list using split()
indicators: [str] = upload.value["indicators.txt"]["content"].decode().split("\n")
indicators

One extra gotcha. See that last empty string in the list? To sort that out, we’ll use the strip() method before split() to remove the trailing newline before turning the string into a list.

# Get the file contents as a list using strip() to remove the trailing newline, then split() 
indicators: [str] = upload.value["indicators.txt"]["content"].decode().strip().split("\n")
indicators

Et voilà! We have a nice list of indicators.

Now, if only we had some way of classifying them…

I know this was a long journey for a single widget, but a lot goes into making this thing work well. In the next lesson, we’ll put this thing to use in making a real tool!

2-4: LAB | File Hash Generator

Now that we have some of these widgets in our toolbelt, we’re ready to make our first real tool: a file hash generator! Is it a little contrived: maybe, but we can build this in such a way that it’ll be pretty useful.

Features

Our generator will be able to:

  • Accept multiple files
  • Show MD5, SHA1, SHA256, and SHA512 hashes all at once
  • Display results in a clean manner

Not so bad, really. Even though there are command line tools that do this, a little user interface ergonomics goes a long way.

Note

This is the first Lab of Part 2. The BEST WAY to use this Notebook is to recreate the tool in your own Notebook. Use this as a guide, but produce the working tool for yourself!

File Hashes in Python

If you’re in this course, there’s a reasonable expectation that you’re familiar with the general concept of file hashes. If you need a refresher, check out this article from SentinelOne.

To make hashes in Python, we have to use the hashlib module. Let’s import it and start playing with it.

import hashlib

hashlib is a very powerful module, but let’s dig down into the sha256 function first.

hashlib.sha256?

Okay, so it returns a “hash object,” whatever that is. And we can give it a string (actually a bytestring) to kick it off. We’ll use the string Python Rocks! for our test.

Before we do the Python version of this, let’s use Jupyter’s shell superpowers to get the Linux command line version for reference:

# Use Jupyter's ability to run shell commands to get the sha256 hash of our test string
! echo -n "Python Rocks!" | sha256sum

Cool. Now let’s try the same thing with Python. Don’t forget, we need to use a bytes object, so watch for that b"" syntax.

# Hashing, Python style
test_hash = hashlib.sha256(b"Python Rocks!")
test_hash

Uh, what.

Okay so we got back the HASH object, but that doesn’t actually tell us anything useful. Let’s inspect the object a bit more.

test_hash?

The two methods that look promising for our purposes are digest() and hexdigest(). If you look closely at the result of the Linux sha256sum command, the characters are all hexademical digits. That should be a hint as to which one we want here.

# Get the proper value from our hash object
test_hash.hexdigest()

Look at that! Our hexdigest() matches! Now that we know how to generate hashes, the next step is to upload some files. I’ve provided 2 sample files here for you to play with.

Upload Files

To make the file uploading easy, let’s provide an upload widget. And of course, import the required modules.

We’re also going to use a nice HBox widget for clean layout of the upload and the label.

# It's importin' time
from IPython.display import display
import ipywidgets as widgets
upload = widgets.FileUpload(multiple=True)
label = widgets.Label(value="Upload sample files here 👉")
hbox = widgets.HBox([label, upload])
display(hbox)

Generate Hashes

With the files uploaded, we can generate hashes for each. There are lot of ways to write this code. Given that we’ll need the same data—md5, sha1, sha256, and sha512 hashes—for each file, my move would be to write a generate_hashes() function. The function will return a dict with hash types as keys, to make our results easy to work with.

I’m going to write this function in a very me style, meaning the code will balance succinctness with readability, but try to minimize intermediate variables. You do not have to write this way.

# Define the `generate_hashes()` function

def generate_hashes(sample: bytes) -> dict:
  """
  Generates a dict of hashes for a given bytes sample
  """
  return {
    "md5": hashlib.md5(sample).hexdigest(),
    "sha1": hashlib.sha1(sample).hexdigest(),
    "sha256": hashlib.sha256(sample).hexdigest(),
    "sha512": hashlib.sha512(sample).hexdigest()
  }

With the function defined and our files uploaded, all we need is some code to grab the hashes. We’ll once more take advantage of nested dictionaries.

Now, we can do this with a for loop, but we can actually expand on a trick we already know. We’ve seen that list comprehensions are a fast way to make a new list from an existing list. Well, turns out there is also a dictionary comprehension that allows us to create new dictionaries from existing ones! Since we have a dictionary from the Upload widget, we can take advantage of this trick to quickly generate a new dict with filenames as top-level keys, and the result of generate_hashes() as the values for each.

To get the right data shape for this, we use the dict’s items() method, which provides a sequence of tuples of shape (key, value) for each item in the dict. We can then use those values in our comprehension.

# Quickly make our file_hashes with a dict comprehension
file_hashes: dict = { key: generate_hashes(value["content"]) for (key,value) in upload.value.items() }
file_hashes

Rad, so we have our hashes! Now, how to display them? Ideally I’d want a nice clean table. Luckily, there is an HTML widget that will allow us to create just such a thing.

As a refresher, HTML tables are structured like so:

<table>
  <tr>
    <th>Column Header 1</th>
    <th>Column Header 2</th>
  </tr>
  <tr>
    <td>Table cell</td>
    <td>Table cell</td>
  </tr>
</table>

We can build up our table with the power of string concatenation. We’ll loop over the keys in file_hashes to get our data.

# Build our HTML table

table = "<table>"
table += "<tr><th>File Name</th><th>MD5</th><th>SHA1</th><th>SHA256</th><th>SHA512</th>"

# Use a loop over file_hashes to access what we need and add to the table
for f in file_hashes:
  table += f"<tr><td>{f}</td>"
  # Use another loop for the hashes
  # Note we access the dict with the key `f`
  for h in file_hashes[f]:
    hash = file_hashes[f][h]
    table += f"<td>{hash}</td>"
  table += "</tr>"
table += "</table>"

Widget time! Let’s display our hard work with an HTML widget:

html = widgets.HTML(value=table)
display(html)

And that’s it! You can of course add some <style> in there to change up the appearance. But you now have a file hash generator that will create multiple hash types for as many files at once as you like!

3-1: Opening Files

It might seem a little silly to cover opening files after going through all that work with widgets. But there are going to be plenty of times when we want to open files without a widget. Widgetlessly. That’s an adverb now.

Plaintext vs. Binary

Most of the files we’ll deal with are plaintext files. Put another way, files we can read with cat or a simple text editor. That’s the kind of file we’ll be focusing on in this lesson. There are plenty of modules that allow us to manipulate more complex filetypes like PDFs or images, but we’ll get to those later.

open() and close()

Python has a built-in function called open() that will access the data at a given path on the filesystem. Open has a few optional arguments, but to simply read a plaintext file, we can simply use open(<path>). You might think this would automatically give us str or bytes data, but nope! This creates a TextIOWrapper object which has additional methods to access its contents. read(), for example, will read in the content as a str.

Once we’ve finished doing whatever we want with the file, we would call the close() method on the IOWrapper to clean up. Check it out.

# Opening and reading files the long way with open() and close()
f = open("indicators.txt")
data:str = f.read()
f.close()
data

The Shortcuts: with and readlines()

If this seems like it’s setting us up for a bunch of extra work…it kinda is. Luckily, Python has a few more tricks to make working with files a bit easier.

The with syntax allows us to open the file and have it automatically closed after we’ve finished with whatever we want in the with block. This structure is for designating temporary resources for a block of code. However, any variables we make inside this block are not temporary.

Another handy method I use is readlines(). Where read() will produce a single string of data from a file, readlines() will produce a list of strings representing the lines of a plaintext file.

# Getting lines the quick way
# Note the with...as syntax
with open("indicators.txt") as f:
  data: [str] = f.readlines()

# And outside of the with block, we still have data
data

You may have noticed that our list items came with the \n attached. That is going to be annoying later on, so my normal process of importing text data adds uses a list comprehension and the strip() method with readlines() to remove those trailing linebreaks.

# Same process, but with a list comprehension and .strip()
with open("indicators.txt") as f:
  data: [str] = [l.strip() for l in f.readlines()]
data

Burn this pattern into your brain. You’ll use it constantly.

Writing Files

We can use the same pattern to write files, but we need to give open() an additional argument: the open mode. By default, this is "r", for read. However, we can also open in "w" for write, or "a" for write-append. What’s the difference? Write will overwrite the contents of a file with new data, while write-append will add on to existing data. You can check out all the open modes in the Python Docs.

Either way, we can use the .write() method of the file object to send string data to the file.

Let’s try it out.

# Write a file, then show the result
msg_1: str = "I'm written first"
msg_2: str = "I'm written second"

with open("example.txt", "w") as f:
  f.write(msg_1)
    
!cat example.txt
# Now let's overwrite it
with open("example.txt", "w") as f:
  f.write(msg_2)
    
!cat example.txt
# Append to a file
with open("example.txt", "a") as f:
  f.write(msg_1)
    
!cat example.txt

Notice that write() did not automatically add a new line on that append. write() is fairly primitive, so if we want our lines to be, y’know, lines, we can either loop through, concatenating our string with "\n", or we can use the .join() trick to join a list of strings. Watch this:

# Join a list of strings with a newline as the separator
"\n".join([msg_1, msg_2])
# Use that .join() trick to write the file
# Append to a file
with open("example.txt", "w") as f:
  f.write("\n".join([msg_1, msg_2]))
    
!cat example.txt
# Cleanup the example file
! rm example.txt

3-2: CSVs

Comma Separated Value files are probably the single most common file format you’ll work with when parsing data. You’ll also encounter a lot of PDFs, but PDFs are not a file format; they are torment engines meant to visit suffering upon the human race. So for now, we’ll stick with CSVs.

Are CSVs that complex? Not really. They’re lines of text with fields separated by commas. We could make do with the tools we’ve already learned, but we don’t have to. Turns out that Python has a fantastic csv module that makes working with this file type incredibly simple.

For this exercise, I’ve provided a CSV of ransomware IOCs, sourced directly from US-CERT #StopRansomware alerts.

Let’s take a quick look with the shell command head.

!head ransomware_iocs.csv

So we have a header row showing us 3 columns: Indicator, Type, and Ransomware Family.

You might be thinking we could use our Indicator class from Part 1 to give these some structure when we import them. And you’re not wrong! But let’s focus on the CSV itself for now.

To work with these files in Python, we’ll begin by importing the csv module. Then we’ll use the dir() function to see what’s inside.

# Import the csv module
import csv

# Let's inspect it
dir(csv)

Reading

There’s a lot going on, but of particular interest* to us are the following classes and functions:

  • reader(): Produces a generic CSV reader from a file object
  • writer(): Produces a generic CSV writer from a file object
  • DictReader(): Produces a reader that uses dictionaries to maintain headers
  • DictWriter(): Produces a writer that writes using dictionaries for headers

Readers are iterators over the rows.

As always, this will make more sense in context. We use the readers and writers in the with block.

Note

Yes I see the excel options. I don’t care because working with Excel files is awful. Plain text 4 Lyfe.

# Get data with a basic reader()

with open("ransomware_iocs.csv") as f:
  iocs = [row for row in csv.reader(f)]

# Check out the first 10
iocs[0:10]

So what did we get? We got ourselves a list of lists. So the csv.reader broke down the file and split each row into a list of values.

Note

The header is included as the first row, so if you’re expect just the data, you’ll want to slice off the first element, like iocs[1:].

Now let’s see how DictReader works.

# Get data with a DictReader()
with open("ransomware_iocs.csv") as f:
  iocs = [row for row in csv.DictReader(f)]

# Check out the first 10
iocs[0:10]

Whoah! So instead of a header row in our list, DictReader detected the header and used it as keys to create a list of dicts!

This is almost always how I import CSVs. The dict provides just enough structure to be useful without needing the whole overhead of a class.

Also, as we’ll see in a later section, lists of keyed-alike dicts play very well with our data analysis library of choice.

Writing

Writing is a very similar process. writer() expects a list of lists. You can either loop through the outer list and call writer.writerow(), or you can do them all at once with writer.writerows(). It depends on if you need to do any modification to the list before you write it.

And DictWriter? Same deal, but it expects a list of keyed-alike dicts.

Let’s create some test data in both shapes to play with.

# Create sandbox data

ship_list: [list] = [
  ["Name", "Registry", "Class"],
  ["Enterprise", "NCC-1701", "Constitution"],
  ["Reliant", "NCC-1864", "Miranda"],
  ["Voyager", "NCC-74656", "Intrepid"],
  ["Discovery", "NCC-1034", "Crossfield"] 
]

# If you think I'm writing that out twice you're crazy
ship_dict: [dict] = [ {"Name": s[0], "Registry": s[1], "Class": s[2] } for s in ship_list[1:] ]

Finally, a dictionary comprehension in action! That’s the kind of Python dorkery I love to do.

Okay, let’s write these out.

# First, with the writer()

with open("ships.csv", "w") as f:
  writer = csv.writer(f)
  writer.writerows(ship_list)
    
# And check the results
! cat ships.csv

Now DictWriter has a bit of a gotcha. It requires a fieldnames named argument. We can extract the fieldnames from any one of our dicts with the .keys() method.

Additionally, if you want the header present, you need to separately call .writeheader() before .writerows().

# Now with DictWriter()

with open("ships.csv", "w") as f:
  # Get fieldnames from the first element
  fieldnames = ship_dict[0].keys()
  # Notice the named arg
  writer = csv.DictWriter(f, fieldnames=fieldnames)
  # Write the header, then the wrows
  writer.writeheader()
  writer.writerows(ship_dict)
    
# And check the results
! cat ships.csv
# Cleanup the folder
! rm ships.csv

No dedicated check for understanding for this section, as we have more file formats to cover! But feel free to play a bit more with CSVs in another Notebook!

3-3: JSON

JavaScript Object Notation is a common format for transmitting data over web APIs. Much like dicts, they are key-value pairs. In fact, Python’s json modules takes advantage of this symmetry to easily read/write JSON data!

Now, instead of recreating a new sample file, why not use the one we already had from the last section? Let’s reuse the csv module to import the IOCs as dicts.

# Produce a handy sample data set with the file from the last lesson
import csv

# Grab the file from the previous section. It's my course; I'll do what I want
with open("../3-2_csvs/ransomware_iocs.csv") as f:
    reader = csv.DictReader(f)
    iocs: [dict] = [i for i in reader]

# Show a sample
iocs[0]

With the IoCs imported as dicts, we can easily make JSON out of them. We’ll start by using json’s dumps function, which produces a string representation of a JSON-able object. Our iocs list absolutely qualifies.

# Use json to JSONify our IoC list
import json

# Here comes a wall of text
json.dumps(iocs)

That’s a lot of text!

What matters here is we now have a well-formatted JSON string. We can write this to a file (but there’s actually an easier way), but we can also send this over network connections—say, to an API endpoint.

Writing Files

The json module makes it even simpler than csv to write files. The dump() method takes a file object and writes the entire JSON-able object to the file.

# Write our IoCs to a file

with open("ransomware_iocs.json", "w") as f:
    json.dump(iocs, f)

Now we have a new .json file in this folder! It’s that simple.

I will also call out that json.dump() has an indent argument that will make the resulting file formatted with indendation and line breaks for human readability. Open this next file up in a text editor to see.

with open("ransomware_iocs_pretty.json", "w") as f:
    json.dump(iocs, f, indent=4)

Quite the difference, right?

Importing JSON

The json module has equivalent reading functions to its writing functions. load() loads JSON data directly from file pointers into Python objects, and loads() will do the same with strings.

# Quick demo of loads
# Note that JSON requires double quotes, so we enclose the str in single quotes
json.loads('[{"foo":"bar"}]')
# Now, let's reload the ransomware IoCs directly from the file

# We start by deleting the iocs variable to remove all doubt
del iocs

with open("ransomware_iocs.json") as f:
    iocs: [dict] = json.load(f)

If it seems like working with JSON is a lot easier than working with CSVs…you might not be wrong. But not all data types can be stored as JSON! Anything not supported by JSON has to be stringified one way or another to be saved as JSON.

Again, no need for a check for understanding yet—we’re saving that all for the final lab for this section 😉

Up next, we take a brief detour into regular expressions!

# Cleanup the files

! rm ransomware*.json

3-4: Regular Expressions

I want to be very clear: this is not an introduction to regular expressions in general. If you need to review the basic concepts, I recommend going through our free regex course first. For a quicker review on syntax, check the Python docs.

Instead, we’ll review how to use regular expressions in Python, and how they become useful in common defensive tasks.

It all begins with the re module, so let’s get that imported.

# Import the re module
import re

Defining Regexes

The first method to learn in the re module is re.compile(), which will take a string and turn it into a Regular Expression object to be used for pattern matching.

This isn’t technically necessary, as we’ll see, but it’s a good practice because it makes the matching process more efficient when reusing the pattern.

We’ll start simple with a pattern for email addresses. What will we need?

An email can be any alpahnumeric, dots, and dashes, and underscores (yes, really, even though GMail prohibits them). The name will be followed by an @ symbol, then a valid domain.

Now since we’re dealing with dots, and dots have a meaning in regexes (any character), we’ll need to escape them with a backslash. But! Backslashes have meanings in strings for escaping characters already!

To avoid this weird gotcha, we can use Python’s raw string syntax to define regexes without any escaping at all. Raw strings are prepending with r. Observe:

# Print backslashes with raw strings
print(r"You'd think this would break a line: \n. And yet, it doesn't in raw mode.")

So we can use raw strings with re.compile() to generate a regex for email addresses. Something a little like:

# Generate the email regex. There are others, but this will do fine for now.
email_pattern = re.compile(r"[a-zA-Z0-9_\-\+\.]+\@[a-zA-Z0-9-]+\.[a-zA-Z0-9-]{2,}")

Matching

With our pattern compiled, we can test this against good (and bad) email addresses to confirm whether it works. To do so, re has several methods that work in subtly different ways.

  • .match() looks for a match against a pattern at the beginning of a string. If there’s a match later, too bad.
  • .search() looks for the first match anywhere in the string. But only the first.
  • .findall() will return a list of all matches in the string.

Which one you want is a question of your performance needs, whether you need all matches, and the shape of your data.

For .match() and .search(), a match object is returned if there were matches. Otherwise, None is returned. This can be useful for conditionals based on matches, as you’ll recall that None is falsy.

To see what was matched, call the .group() method.

Let’s demo the results of both with a sample of multiple good and bad emails.

# Build our sample
sample: str = "test@test.com testattest.com test@test.c test@bad_domain.com tes+t_1@good-domain.com"

# Test with match
match = email_pattern.match(sample)
# Review the findings
match.group()
# => 'test@test.com'

So obviously we get one match, but more than one of these should work. Let’s try with .findall()

# Now try with .findall()
all_matches = re.findall(email_pattern, sample)
all_matches
# => ['test@test.com', 'tes+t_1@good-domain.com']

There we go!

Named Groups

One of my favorite tricks with regular expressions is using named groups. With this regex feature, we get match groups that we can reference using 3 methods: .group(), .groups(), and .groupdict(). The .group() method allows one-at-a-time access to named groups; the .groups() method returns a tuple of all matches; .groupdict() returns a dict with the named groups as keys. This in particular can make really quick work of parsing files such as logs.

The Notebook folder has an auth.log file for us to play with. Let’s import it as lines to process further.

# Import auth.log
with open("auth.log") as f:
  auth_logs = [l.strip() for l in f.readlines()]

# Display them
auth_logs

"""
=>
['10 Oct 2022 15:22:31 - User admin logged in',
 '10 Oct 2022 15:25:10 - User bob logged in',
 '10 Oct 2022 16:02:24 - User admin logged out']
"""

If we examine the log—simplistic as it is—and consider what we want out of it, 3 items emerge:

  • A timestamp
  • A user
  • An action

Since each line has a predictable pattern, we can write a regex with named groups to capture the data and save the results. Let’s build the pattern.

  • The timestamp is the beginning of the line, followed by a 2-digit number, a space, a three-character capitalized month, a 4 digit year, and then a HH:MM:SS time
  • The user is after the string User and followed by a space
  • The action comes after the space following the user.

Capture groups in regexes are delimited by parentheses. In Python, we make named groups with the (?<name>pattern) syntax.

Putting it all together now…

log_pattern = re.compile(r"^(?P<timestamp>\d{2} [A-Z][a-z]{2} \d{4} \d{2}:\d{2}:\d{2}) - User (?P<user>[a-z]+) (?P<action>.+)$")

Note

It took me years to be this fluent in regular expressions, and I still make mistakes all the time. If the above is difficult to parse, don’t sweat it. Try to break it down one section at a time. I’ve also created a breakdown of this pattern on regex101, which will explain each section and show the matches.

Let’s see what we get with a match!

# Test match
match = log_pattern.match(auth_logs[0])
match

"""
=>
<re.Match object; span=(0, 43), match='10 Oct 2022 15:22:31 - User admin logged in'>
"""

Now let’s see what we get across our three group methods.

# Check .group()
match.group("timestamp")
# => '10 Oct 2022 15:22:31'
# Check .groups()
match.groups()
# => ('10 Oct 2022 15:22:31', 'admin', 'logged in')
# Check .groupdict()
match.groupdict()
# => {'timestamp': '10 Oct 2022 15:22:31', 'user': 'admin', 'action': 'logged in'}

Plainly, we have a successful match! Since our pattern works, we can use it in a list comprehension to quickly parse all the logs.

logs_parsed: [dict] = [log_pattern.match(l).groupdict() for l in auth_logs]
logs_parsed
"""
=>
[{'timestamp': '10 Oct 2022 15:22:31', 'user': 'admin', 'action': 'logged in'},
 {'timestamp': '10 Oct 2022 15:25:10', 'user': 'bob', 'action': 'logged in'},
 {'timestamp': '10 Oct 2022 16:02:24',
  'user': 'admin',
  'action': 'logged out'}]
"""

Success! We now have a list of parsed logs in dicts.

This method of parsing otherwise difficult-to-parse logs has been invaluable in my defensive work. Another example is this Apache HTTP log parser from a collection of defensive notebooks I’ve been cobbling together. There’s a lot going on in there which we’ll cover shortly.

For now, let’s put all this together in our closing lab for this unit!

3-5: LAB | IoC Extractor

It’s time to build a real tool with all we’ve learned! We’re going to solve a GIANT problem defenders face constantly: extracting indicators from garbage source files.

Sometimes, indicators are received in mixed formats. MD5 hashes, SHA1s, SHA256s, IPs, domains, and URLs are all jumbled together.

Sometimes, some sadistic monster decides to provide all of these in a PDF for no reason other than to make us suffer.

But that ends now.

Our IoC extractor will take any text file, find all the strings matching our indicator patterns, and produce either a CSV or JSON for easy ingestion into our defensive tools.

We’ll build it using widgets to make it easy to work with.

Let’s import what we’ll need.

NOTE: We’re gonna use a new library, pdfplumber, to handle PDFs. Don’t worry; we’ll cover it when it comes up.

# Get our gear
import ipywidgets as widgets
from IPython.display import display
import json
import csv
import pdfplumber
import re

Step 1: Acquire Files

First we need the files. Let’s get ’em, widget-style!

# Define and display widgets
upload = widgets.FileUpload(multiple=True, accept=".csv, .txt, .json, .pdf, .html")
label = widgets.Label(value="Upload File(s)")
box = widgets.HBox([label, upload])
display(box)

Step 2: Extract Raw Content

Something kind of counterintuitive about this process is that, with the exception of PDFs, we don’t really care about what kind of file it us. We just want the raw text, because we’ll be doing our own parsing of the content to look for our patterns. PDFs present a special challenge because they are not plain text; they are binary blobs. That’s why we imported the pdfplumber module to help us out.

Let’s briefly discuss how to use pdfplumber.

pdfplumber

Since PDFs are binary blobs, we can’t just do with...open...readlines(). But that’s okay; the plumber has us covered. pdfplumber.open() takes a path and returns a PDF object. This object in turn contains a pages property that holds a bunch of Page objects. FINALLY, those Page objects have an .extract_text() method that will stringify the contents of the page! Check it out.

# Extract text
with pdfplumber.open("esxi_mandiant.pdf") as pdf:
  text_pages = [p.extract_text() for p in pdf.pages]
  
# Review page 1
text_pages[0]

Okay, so we plainly have a way to handle PDFs. Now we need the entire procedure for handling all our files. If it’s plaintext, we can just grab it from the widget. But if it’s binary, we’d be better off using pdfplumber. My first move would be to define some functions.

# First our PDF extractor
def get_pdf_text(pdf_path: str) -> str:
  """
  Extracts text from file at pdf_path and returns a big ol' string of the results
  """
  with pdfplumber.open(pdf_path) as pdf:
    return "".join([p.extract_text() for p in pdf.pages])
  
def get_file_contents(index: int, uploads: tuple) -> str:
  """
  Seeks the upload widget value for a given filename.
  
  If it's there and it's not a PDF, grabs the content as a string.
  
  PDFs, it will use the filename with get_pdf_text
  """
  # Check for a PDF
  if uploads[index]["type"] == "application/pdf":
    return get_pdf_text(uploads[index]["name"])
  
  # Otherwise, get the contents
  return bytes(uploads[index]["content"]).decode()
  

Okay! Functions in hand, we can go get our data. We’ll keep everything in a dict so we’ll once again use that dict comprehension trick we saw before.

# Get data
data: dict = {f['name']: get_file_contents(i, upload.value) for i, f in enumerate(upload.value)}

You can look at data if you want, but it’s gonna be a hot mess so I’m leaving it alone for right now.

Step 3: Match Patterns

Data in hand, we need some ✨regular expressions✨ to match our patterns. What do our indicators look like? Luckily, our hashes have known lengths and URLs/domains/IPs are fairly simple.

  • MD5: 32 characters of lower case 0-9,a-f
  • SHA1: 40 characters of the same
  • SHA256: 64 “”
  • SHA512: 128 “”
  • IPv4: 1-3 digits and a dot, repeated 3 times, followed by 1-3 more digits
  • domain: A-Z,a-z,0-9,- separated by dots, ending in 2-n a-z characters
  • URL: http, maybe s ://, a domain pattern, and any number of / followed by the same characters allowed in domains,plus%,?,=,&, and +.

Now, we do have an extra complication here: overlaps. A SHA512 will contain 4 MD5 pattern matches! How do we avoid this issue? Negative lookarounds. We will use the (?<![0-9a-f]) and (?![0-9a-f]) patterns to discard matches with hex characters immediately before or after the smaller matches. This is some advanced regexery, but it’s what we need here.

It’s Regex time!

# Define our regexes
# Yes, I'm giving these to you. Feel free to improve on them!
md5_pattern = re.compile(r"(?<![0-9a-f])[0-9a-f]{32}(?![0-9a-f])")
sha1_pattern = re.compile(r"(?<![0-9a-f])[0-9a-f]{40}(?![0-9a-f])")
sha256_pattern = re.compile(r"(?<![0-9a-f])[0-9a-f]{64}(?![0-9a-f])")
sha512_pattern = re.compile(r"[0-9a-f]{128}")
ipv4_pattern = re.compile(r"(?:[0-9]{1,3}\.){3}[0-9]{1,3}")
domain_pattern = re.compile(r"(?:[A-Za-z0-9\-]+\.)+[A-Za-z]{2,}")
url_pattern = re.compile(r"https?://(?:[A-Za-z0-9\-]+\.)+[A-Za-z0-9]{2,}(?::\d{1,5})?[/A-Za-z0-9\-%?=\+\.]+")

Yeesh. With those in hand, it’s time to parse our data. Again, we’re building a dict whose keys are our filenames. Inside each is a dict whose keys are our match types. Easiest way here is with a loop.

You might notice an unfamiliar bit of syntax in those regexes: (?:...). These are non-capturing groups which are necessary to prevent .findall() from returning just that group.

Oh, one other thing. It’s likely that a given file will have repeated indicators. We need an easy way to deduplicate our findings. We could write an algorithm to check and remove duplicates, but there’s an easier way.

The set data type is a sequence that enforces uniqueness. Converting a list to a set automatically deduplicates the elements. Sets are immutable, so we lose some flexibility, but they’re fantastic for this exact use case.

But, because we can’t save sets as JSON, we’re going to immediately convert them back to lists. This is the kind of nonsense that you’ll see often in my code, and it’s a symptom of training as a functional programmer. There are a lot of parentheses.

# Initialize an empty dict
results = {}

# Loop through our strings
# Convert our findings to sets to deduplicate
for d in data:
  content = data[d]
  results[d] = {
    "md5": list(set(md5_pattern.findall(content))),
    "sha1": list(set(sha1_pattern.findall(content))),
    "sha256": list(set(sha256_pattern.findall(content))),
    "sha512": list(set(sha512_pattern.findall(content))),
    "ipv4": list(set(ipv4_pattern.findall(content))),
    "domain": list(set(domain_pattern.findall(content))),
    "url": list(set(url_pattern.findall(content))),
  }
results

Okay, so is it perfect? No. It turns out files with extensions are basically indistinguishable from domains. But otherwise, pretty solid! Let’s take this and deliver our results. To do so, we’ll offer 2 buttons: a CSV export, and a JSON export.

Buttons are event-driven, so we also need to first write the functions that will perform the export.

Step 4: Deliver Results

Reshaping our data into JSON format is essentially done for us. But flattening it out into a CSV, well, that’s gonna be a little trickier.

If we imagine our CSV with 3 columns, like so:

FilenameIOC TypeValue
Foo.txtdomainevil.com

We can see that a bit of destructuring is necessary.

For now, rather than be super elegant with it, we’ll be clear. Sadly, this means it’ll involve 2 nested loops (gross) and a list comprehension. So really 3 nested loops, but one is prettier.

# Give an output
output = widgets.Output()
# Create the buttons.
json_button = widgets.Button(description="JSON Export")
csv_button = widgets.Button(description="CSV Export")

# Provide an output display
box = widgets.HBox([csv_button, json_button, output])
display(box)

# Define filenames
CSV_RESULTS: str = "results.csv"
JSON_RESULTS: str = "results.json"

# Create the export handlers
def csv_export(b):
  
  with output:
    header = ["filename", "type","value"]
  with open(CSV_RESULTS, "w") as f:
    writer = csv.writer(f)
    writer.writerow(header)
    for filename in results:
      # Add filename to the the dict
      result = results[filename]
      for ioc_type in result:                    
        iocs = result[ioc_type]
        rows = [[filename, ioc_type, i] for i in iocs]
        writer.writerows(rows)
                
    # Notify when done
    print(f"{CSV_RESULTS} written")
      
def json_export(b):
  
  with output:
    with open(JSON_RESULTS, "w") as f:
      json.dump(results, f)
      print(f"{JSON_RESULTS} written")


csv_button.on_click(csv_export)
json_button.on_click(json_export)

Step 5: Celebrate!

That’s it! We now have a take-all-files IoC extractor. Remember that this file is meant to explain the lab, but that you should try to reproduce it in your own Notebook to understand in depth how each piece works.

At long last, we’re finished with our unit on Parsing. Up next, we move off our one computer and use Jupyter to connect to resources on the internet!

References

  1. Mandiant, 29 September 2022. “Bad VIB(E)s Part One: Investigating Novel Malware Persistence Within ESXi Hypervisors”
  2. AlienVault, 12 September 2022. “HEINEKEN Malaysia ‘Free Beer’ phishing scam circulating through social networks which has been detected by NetAssist Threat Intelligence team”
  3. CISA, 30 June 2022. “#StopRansomware: MedusaLocker”

4-1: Requests

We’re about to go on a magical journey. A journey outside of space and time, into a realm I like to call…

The information superhighway.

Yes, we’ve all navigated the internet a time or two—usually with a web browser! But it turns out Python has the ability to collect information from the internet as well, and we’re going to use that ability to connect our Notebooks to powerful information sources.

To do, we must harness the power of HTTP. Luckily, Python has a handy library to do so: requests

# Import my bff, requests
import requests
# A simple requests demo
url: str = "https://taggart-tech.com"
r = requests.get(url)
r.headers

Look at that! With a simple command we were able to reach out and get the headers (and more) from my website!

requests is incredibly powerful. With it we can submit form data, connect to APIs, and even handle JSON. It underpins most of the API-based libraries that defenders may use like VirusTotal and Shodan.

We’ve already seen how to use the library to make HTTP GET requests. As you might imagine, this is just the tip of the iceberg. I would get real comfy with the Requests Documentation now. I’m constantly referencing it to make sure I’m going things correctly.

For now, let’s dir to examine a Response object returned by requests.get().

dir(r)
# => Looooong

Of particular interest will be the status_code, which will inform us whether a request was successful, and the text, content, or json properties, depending on the type of data returned. text contains str of the response body; content contains a bytes of the same; and if the data can be parsed as JSON, json already has the dict representation waiting for us!

4-2: APIs

Application Programming Interfaces are ways for one piece of software to communicate with another. In our case, we are specifically referring to HTTP endpoints where we can retrieve and send data. These are often referred to as REST APIs.

The potential value of connecting our Notebooks to external data sources cannot be overstated. By leveraging the power of the interwebs, we can verify our findings, submit new intelligence, and enrich/correlate our data to tell a more complete story.

To demonstrate this power, we will play a bit with the VirusTotal API. If you’re not familiar with VT, it’s a fantastic resources for community-submitted malware samples and indicators. It’s a primary part of my workflow.

In order to do this, you will need a VirusTotal API Key, which means you’ll need a VirusTotal account. They are free! Go get one. I’ll wait.

…

Got one? Cool. Now, don’t tell anyone. Not even me.

Seeeecrets

Okay, got your API Key? Cool. One thing that we’ll commonly have to do is enter secrets like API keys or passwords into a Notebook to handle authentication. This is tricky because we certainly don’t want to expose those in plaintext, and we definitely don’t want to save them in the state of our Notebook unencrypted. What’s a Jupyteer (I just made that up) to do?

The getpass module allows us to collect secrets in a Jupyter Notebook and use them securely. They are not stored in the saved state of the Notebook, meaning this process it Git-safe!

Let’s import it and play with it.

from getpass import getpass
secret = getpass("Gimme a secret!")

Now you can print that out, but don’t. Just use these secrets as you need them.

VT The Hard Way

VirusTotal does have a Python module we can use, but I want to show you how to use the REST API directly first.

Headers and Authentication

It is very common for REST APIs to use a HTTP header for authentication. In VT’s case, the header x-apikey must be present and set to a valid API key. Let’s set that up now (and import requests)

# Import our stuff
import requests
import json
# Set the API key and build the headers
api_key = getpass("VT API KEY")
headers = {
  "content-type": "application/json",
  "x-apikey": api_key
}

Now we need a sample to test. I have one for us: a sha256sum of a strange file I found on a webserver:

SHA256: 1c263b3f4d21039b2a89865a4ab6600f1cc034817bae6ab1f91599674e94be72

Now let’s build our API request. Luckily, the endpoint for Files is quite simple: a GET request to https://www.virustotal.com/api/v3/files/[FILE HASH].

With that and our header, we should be good! Let’s build that with requests and get our data!

# Build the VT Search

# How else could we get hashes? Think back to what we've done before.
search_item: str = "1c263b3f4d21039b2a89865a4ab6600f1cc034817bae6ab1f91599674e94be72"

url: str = f"https://www.virustotal.com/api/v3/files/{search_item}"

# Note the use of headers in the GET requests
# We want the JSON result
res: dict = requests.get(url, headers=headers).json()

The thing about VirusTotal responses is that they can be kind of enormous. Let’s break this one down to see what we have.

res.keys()

Okay, 2 keys. Not so bad right? Well…data is a little bigger.

data: dict = res["data"]
data.keys()
# Now let's see attributes
attributes: dict = data["attributes"]
attributes.keys()
# => Many keys

Obviously there’s a lot to look through. It helps to know what we’re after. I like type_description, names and popular_threat_classification for starters. These help identify what kind of a thing it probably is.

# Print some basic data
print("Type Description:")
print(attributes["type_description"])
print("\nNames:")
print(attributes["names"])
print("\nPopular Threat Classification:")
print(attributes["popular_threat_classification"])

If you want to go a bit deeper on the results, sandbox_verdicts is always interesting.

attributes["sandbox_verdicts"]

And of course if you want the whole list of detections, that’ll be in last_analysis_results. Because the list is so huge, I might deconstruct it a little bit. Also, knowing that failed detections are Nones, I might use that to see just the engines that did detect it.

# Use a pretty complex list comprehension to get just successful detections
positive_detections: [tuple] = [(k, v["result"]) for k, v in attributes["last_analysis_results"].items() if v["result"]]
positive_detections

At this point, we have a pretty good idea that our hash belongs to a Linux executable that is a Mirai variant. If this came from a machine under our purview, we’d have cause for alarm! Or at the very least, incident response.

VT The Easy Way

Now that we’ve explored using the API “raw,” we can talk about using the Python module. Now, you don’t have to use it! If you prefer manual HTTP requests, that’s fine. But the library can make some of the ergonomics a little better. Let’s import it and see.

Note

We’re also importing asyncio for reasons explained shortly.

# import the VT library
import vt
import asyncio

Many of these API libraries are built around a Client class, that we instantiate with our credentials. This one does too, but here we must take a quick detour into…

Asynchronous Python

Sometimes, processes are not finished instantaneously. Network requests often take time to complete. If we hold up the works until the request is done, that is considered a synchronous and blocking operation—nothing else can happen until that procedure is finished. However, we can also write asynchronous or non-blocking code, that will come back and do a thing once the non-instant procedure is complete.

VirusTotal’s native Python library is asynchronous. That means we need to use it as such in our Notebook. Jupyter supports asynchronous code; we just need to wrap our async calls in an async with block. Also, we use await to capture the asynchronous output inside of that block. Check out Python’s asyncio documentation for more details.

We’ll define a new client in the with block, passing in our API key. Then we can now perform the same search we did with requests by specifying the path and using get_json() or get_object(). In both cases, you still need the last part of the API path.

# Async time!
async with vt.Client(api_key) as client:
    r = await client.get_object_async(f"/files/{search_item}")
    print(r)

r may not look like much, but it has everything we need easily accessible. Look!

r.type_description
# => ELF
r.sandbox_verdicts

Handy, right? That is much easier than navigating a ton of dicts

4-3: Scraping

APIs are great, but they can’t do everything. Sometimes we need to use other methods to analyze content. Specifically, I want to discuss how we analyze web content either for disposition or to extract useful information from it. Put another way, how can we make Python read and understand websites programmatically?

The answer is BeautifulSoup.

Funny Name, Invaluable Library

This may be one of the most important Python modules I’ve ever used. BeautifulSoup is a parser for HTML/XML files. With it, we can do some pretty incredible things. Let’s start by learning the basic shape of BS4.

We begin by importing the library

# import and setup
import requests
from bs4 import BeautifulSoup

Now we need to give it something to parse. Once again, My simple blog serves well for this.

document: str = requests.get("https://taggart-tech.com").text
soup = BeautifulSoup(document)

soup is a complex object, but it helps us parse the HTML in valuable ways. Let’s grab the page title as an example.

# Access the title
soup.title

At first glance, that might seem like a string of HTML. But look carefully: no quotes. No, this is something else.

# What are you, soup.title?
type(soup.title)

As a Tag object, there’s more we can do with the title. Including extract its text.

soup.title.text

Now that’s a string.

Seeking through HTML

The power of BeautifulSoup comes from the ability to slice through a document for exactly what we want. For example, say we watned to pull all the links (<a>) tags out of the page. We could use the .find_all() method to do just that.

# Grab links
links = soup.find_all("a")
# Show the first one
links[0]

Again, since the results are Tag objects, we can destructure this a little bit more to access just the href attribute.

# Get just the hrefs 
# And use (set) to dedup
hrefs = set([ l["href"] for l in links ])
hrefs

The syntax of all the ways to work with BeautifulSoup objects can be a bit confusing. I almost always have the documentation up while I’m working with it. Still, the capabilities are quite impressive.

An RSS Reader?! Yes.

We can heal the wounds. We can create an RSS Feed processor in Python.

Did you know CISA’s US-CERT has an RSS Feed? It’s true! It’s a solid way to get basic information about emerging threats.

So let’s build handler and parser to:

  1. Process the Feed
  2. Produce Headlines
  3. Get data from linked articles.

To start, let’s set up the URL of the RSS feed:

# Set feed URL
feed_url: str = "https://www.cisa.gov/uscert/ncas/current-activity.xml"

Then we’ll download the data with requests and use BeautifulSoup to parse the XML.

# Get the XML
us_cert_xml = requests.get(feed_url).text
# Soupify the XML

soup = BeautifulSoup(us_cert_xml, "xml")

Now that we have the parsed XML, we need to understand the structure. Every article is inside of an <item> tag. So to get all the items, we will use the .find_all() method.

# Get all <item>s
items = soup.find_all("item")

# Review one for structure
items[0]

So each item has a title, a link, and a description. That’s enough to get us going. But how should we handle it?

One option might be the HTML Widget. We can convert each item into some crafted HTML.

Did you notice the weird characters in description? The HTML tags have been escaped to make them XML safe. To reverse that, we’ll need the html.unescape() method. A simple import, but worth mentioning.

Okay, let’s write a function to HTMLify the items.

from html import unescape

def html_item(item) -> str:
  """
  Converts an XML item into HTML
  """
  item_html = f"<a href=\"{item.link.text}\"><h2>{item.title.text}</h2></a>"
  item_html += f"<p>{unescape(item.description.text)}</p>"
  return item_html

And now, to display the results. I’m adding a little <style> to the proceedings to make the links pop.

import ipywidgets as widgets
from IPython.display import display

style_html = """
<style>
  a {
    color: #9580ff;
    font-weight: boldest;
  }
</style>
"""

# We'll just do the first 3 items 
html_items = [ html_item(i) for i in items[:3] ]
html_str = "".join(html_items)

html_out = widgets.HTML(value=style_html + html_str)
display(html_out)

Now this is one way to parse and use this data, but I’m sure you can come up with others—and other feeds to connect to. Get creative!

In our final lesson in this section, we’ll build a webpage analyzer to assess for potential malware!

4-4: Tracker Detection

It isn’t all about malware, you know.

Being able to safely determine what a website does without using a browser is a common function of a defense team. Of particular interest is any tracking technologies in use by a site. You may have seen some recent news regarding one of the most insidious: the Meta Pixel Another common tracker is, of course, Google Analytics.

If we want to identify whether a site is using these, we can use our parsing powers to do just that.

We’ll start by importing the stuff we know we’ll need.

# Import the basics
from bs4 import BeautifulSoup
import requests
import re

Recon

First, we need to know a little about what we’re trying to identify. These trackers will have identifiable code snippets in page we analyze. For flexibility, we’ll store these patterns in a dict. You can learn about the code that makes these trackers up in their documentation. For both Meta and Google, we’re talking about JavaScript.

Meta uses a script with the obvious pattern of a URL: https://connect.facebook.net/en_US/fbevents.js. Google uses a URL as well, but as a src attribute for its script. Another pattern we can use in the text of a script node is gtag('config'.

So let’s store those for later reference.

# Set up patterns
# Using a raw string to escape the paren
script_patterns = {
  "meta": "https://connect.facebook.net/en_US/fbevents.js",
  "google": r"gtag\('config'",
  "plausible": "plausible.init"
}

Now we need a list of pages to search. You can add whatever pages you like to this list, but I’ll kick it off with a few.

# Build the list of sites
sites = [
  "taggart-tech.com",
  "npr.org",
  "penny-arcade.com",
  "tomshardware.com",
  "meta.com",
  "cnn.com"
]

Our results will be stored in a dict as well. For each site, we need to:

  1. Pull the web content
  2. Find all script tags
  3. Check if any of them match our patterns
  4. Report results

A few key words in these steps give us hints about how to proceed. any() is a built-in. Python function that take 2 arguments: a function, and a sequence over which to iterate the function. The function must return a bool. If any of the elements return True when passed to the function, any() returns True. There’s also an all() that works similarly, but any() will work for our purposes. So let’s write the pattern_match() function.

def pattern_match(pattern: str, sample: str) -> bool:
  """
  Checks whether a pattern is found in a sample
  """
  return re.search(pattern, sample) != None

It isn’t much, but it will provide what we need to pass to any()…sorta.

But first, let’s initialize our results dict. The shape will be {site: [tracker names]}.

results = {}

So now we need to loop through our sites, grab the script elements, and check for patterns with any(). The idea here is that if any of the script elements match our pattern, we know we have the given tracker.

This is one of those instances where a for loop is absolutely more appropriate than a list comprehension for readability and maintainability.

# Loop through sites
for s in sites:
  # Initialize site list
  results[s] = []
  print(f"Fetching {s}")
  # Get content
  content: str = requests.get(f"https://{s}").text
  soup = BeautifulSoup(content)
  # Find script tags
  # Extract just the text to get the raw str
  script_tags = [ t.text for t in soup.find_all("script") ]
  # Loop through patterns to check
  for p in script_patterns:
    pattern: str = script_patterns[p]
    # Check if any of the scripts have the pattern
    if any([pattern_match(pattern, t) for t in script_tags]):
      results[s].append(p)
# Show results
results

Note

Change the sites up for different results!

Going Further

One of my favorite tools for web app testing is Wappalyzer, which will detect technologies in use on a page via server headers and source analysis. We’re going a simplistic version of this here, but wouldn’t you know it: Wappalyzer has a Python module!

Because I’m cool like that, I’ve already added it as a dependency in this environment. See if you can rework this lab using Wappalyzer to detect WordPress sites or other technology!

And that closes out section 4! Up next, we begin to analyze our data with scientific tools!

5-1: Pandas

No, not the useless overgrown rodents. Pandas is an incredibly powerful data science library that allows deep statistical analysis on datasets through an easy-to-use API.

The simplest way I can explain it is this: imagine if Excel had a command line interface.

Tabular Data

The reason I chose Excel as our metaphor is that Pandas deals with tabular data—data represented in rows and columns, like a spreadsheet. It comes packed with tools to parse, filter, summarize, and analyze these rows and columns, far beyond what would be easily accomplishable in a GUI.

To get started, we will need ourselves a dataset. Let’s take this opportunity to solve another common problem in defense: rapid IP/domain name lookups.

Looking up IP/Domain data

Very often, I will need to pull DNS records for a suspicious domain, then immediately pivot on that IP address to a RDAP/whois lookup on that IP address from ARIN data. I’ve gotten pretty quick at doing this on the command line, but especially if I have ot do it for multiple domains/IPs at once, it is nice to have a tool that automates the retrieval, and tabulates the data for me for simple review.

For this, we will leverage three different Python modules: pydig, ipwhois, and geoip2fast. pydig gets us domain → IP conversion; ipwhois looks up the ARIN data; and geoip2fast enriches that data with more accurate geographic locations for subnet.

Our procedure will look like so:

  1. Create the list of IPs/Domains to review
  2. For each entry, do the following:
    1. If it’s a domain name, look up the IP via dig.
    2. With the IP, perform an RDAP lookup on the IP
    3. Collect CIDR, AS name, number, and listed location.

Our final dict for each domain/IP will look like:

{
    ip: str
    domain: str
    asn: int
    as_name: str
    cidr: str
    country: str
    city: str
}

Let’s build our list of domains and then set up the information collection.

# Import our stuff
# Pandas is conventionally named pd
import pandas as pd
import pydig
from ipwhois import IPWhois
from geoip2fast import GeoIP2Fast
from ipaddress import IPv4Address

We need to use geoip2fast’s command line to update its local database. You can do this directly in Bash or in the Jupyter Notebook with:

# We also need to use the geoip2fast command line to install its database
# Omit the ! for direct Bash usage
! geoip2fast --update-all

I’ll use the list of sites from the last lesson, but feel free to modify it! It won’t affect the run of this Notebook.

# Our sites to analyze
sites = [
  "taggart-tech.com",
  "npr.org",
  "penny-arcade.com",
  "tomshardware.com",
  "meta.com",
  "cnn.com"
]

Putting the Data Together

Handling all these API calls and compiling the data is a tricky feat. First, we have to handle the possibility that the type of DNS record we’re looking for isn’t what a given domain name resolves to. A records will return IP addresses, but CNAMEs will return other domains. We have to handle that. We’ll do a best effort depth-1 resolution of CNAMEs to IPs by catching the first error and retrying for CNAME records, then querying that for an A record.

And if even that doesn’t work, we’ll fail over to blank ASN data. By the time we get to building the dict, we should have all the errors handled.

Tip

This is a straight script, but perhaps you can think about how to build it as a function?

# Initialize sites_data to an empty list
sites_data = []

# Initialize the GeoIPFast client
gip = GeoIP2Fast()

for s in sites:
  try:
    dig_res = pydig.query(s, "A") # Querying A records for IP addresses
  except ValueError:
    # If no A records, look for CNAME
    cname_res = pydig.query(s, "CNAME")
    dig_res = pydig.query(cname_res[0], "A")
  resolution = dig_res[0]
  try: 
    addr = IPv4Address(resolution)
    whois_target = ipwhois.IPWhois(resolution) # We'll get the first result
    whois_res = whois_target.lookup_rdap(depth=1)
  except:
    # We create a mock whois_res to satisfy object creation
    whois_res = {"asn": None, "asn_description": None, "asn_cidr": None}
  
  gip_res = gip.lookup(resolution)
  sites_data.append({
    "resolution": resolution,
    "domain": s,
    "asn": whois_res["asn"],
    "as_name": whois_res["asn_description"],
    "cidr": whois_res["asn_cidr"],
    "country": gip_res.country_code,
    "city": gip_res.city.name
  })

Phew! And after all that, what does our data look like?

sites_data

Well that’s readable, but not very. But the way we created this dataset is intentional. Every dict has the same keys and value types, which means it is really easy for us to create a Pandas DataFrame from it. This will produce a data structure that is not only easy to work with, but displays cleanly in Jupyter.

Let’s make a new DataFrame from sites_data. By convention, primary DataFrame variables are named df.

# Create our new DataFrame
df = pd.DataFrame(sites_data)

Exploring DataFrames

DataFrames are a deep topic, and this lesson (and the next) really won’t cover more than a small fraction of what’s possible. If you want to become extremely fluent in DataFrame usage, I encourage you to spend some time reading the Pandas documentation.

But we can still learn some basics. Let’s preview our data with the .head() method, which show the top 5 (by default) values from the DataFrame.

# Look at the top of it with the built-in `.head()` method
df.head()

Looks great, right? In addition to a clean table display, the DataFrame API is powerful. Let’s play a little bit with it. Let’s select just two columns: as_name and cidr.

Note

Pandas syntax looks pretty strange at first, but you’ll get used to it. Here we’re providing the column names as a list to the DataFrame in index notation.

df[["as_name", "cidr"]]

We can also use the .query() method to write structured queries against the data. Let’s look for Fastly in the as_name.

df[df.as_name.str.contains("FASTLY")]

There is so much to explore in DataFrames, and we’re going to dive in for real in the next lesson. This was meant as a quick intro to the concept of using Pandas and how to consider shaping our data to be useful in tables.

See you in the next lesson where we’ll get serious about Pandas! But before we get out of here, I’ll leave you with probably the most import Pandas feature: .to_csv()

df.to_csv("sites_data.csv")

5-2: DataFrames and Series

Now we’re ready for prime time. Working with data in Pandas is some of the most important material in this course. Once you know how to do it, it will transform your analysis capabilities. Do not skip this section; do not skimp on this section. It’s far too valuable for anything less than your complete attention.

Now then, let’s begin.

You’ve already seen the creation of DataFrames from list[dict]-shaped data. Let’s explore some other common uses.

DataFrames from CSVs

This particular pipeline is probably the most common ingestion method to get data into a DataFrame. CSV exports of logs, IoCs, etc. are made extra powerful inside of Pandas. Luckily, Pandas has a built-in method called .read_csv() that will take a CSV and create a DataFrame using a header row for column names.

The CSV we’re going to use is a list of domains associated with the Blackfile/Helix credential stealer campaign. This forms a solid basis for analysis and enrichment.

Let’s load that in now and use .head() to check it out.

# Import Pandas
import pandas as pd
# Use .read_csv() to create our DataFrame
# This may gnerate a DTypeWarning. Don't sweat it.
df = pd.read_csv("blackfile.csv")
# Look at the first 5 rows with .head()
df.head()

What’s in a DataFrame?

A DataFrame contains multitudes. Understanding how each part functions will make manipulating them to do our bidding much, much easier.

Indices

If you look closely at the output of df.head() above, you’ll see that just to the left of the uuid column, there appears to be an extra column. What gives?! We didn’t define that!

Indeed we did not. Every DataFrame requires an Index, which is how the individual rows in the DataFrame are referenced. Any column containing unique values can be an Index, but ideally it’s one containing sensible sequential values. We can define and index column when we create the DataFrame, but if we don’t, Pandas will create one for us. Let’s look at df.index to see what it is.

df.index

Indices are important because they are the way we access rows. It’s important to not think of indices in DataFrames the same way we think of them in a list. They don’t really work the same way. Watch what happens when we try to use the list indexing syntax with a DataFrame:

# Try to get the first row...or will we?
df[0]

Yeah, that doesn’t work. Pandas DataFrames have their own unique syntax for accessing data in this manner, built on two properties: .loc and .iloc. They look a little strange in practice.

Let’s try .loc first.

# Access the first row
df.loc[0]

But that’s not all .loc can do. We can pass column names as well after a comma! Let’s say we wanted the value column.

df.loc[0, "value"]

That’s all we get! What if we want more than one col? List ’em.

df.loc[0, ["uuid", "Description"]]

Okay, but what if we want multiple rows? We can either list them or use a slice syntax.

# Access the first 10 rows
# We can also use a list of index values
df.loc[0:9, ["value", "comment"]]

Notice that we get a much nicer output when we select multiple rows.

.iloc works similarly, except that it uses integer values for access rather than names. In the case of our Index, they are one and the same, but watch how we can access columns with .iloc.

# Using .iloc to access data
df.iloc[0:10, 0:3]

In truth, I rarely access data directly in this way. The whole point is to manipulate the data at scale, so it’s uncommon for me to need to access specific rows.

But before we depart Indices, I want to stress the value of using a custom Index rather than the default. Depending on your data shape, this can be incredibly convenient. One common example is if your dataset has (completely) unique timestamps. In that case, you can convert the timestamp to a DateTimeIndex and have Pandas automatically sort your data chronologically. This also gives you the power to group data by hour, month, day, etc.

Our data happens to have unique values in the value column. So if we wanted to, we could use that as our index. Let’s try it and see the difference.

We can change the index of a DataFrame with .set_index(). I want you to look at the docs for this method because it introduces a common pattern in Pandas: normally, when we make a global change to a DataFrame, Pandas will return a new DataFrame instead of mutating the original. This is for data integrity and is a good idea! However, if you know for sure you want to change the original, you can often pass inplace=True as an optional argument to the method.

But we won’t be doing that right now.

# Create domain-based index
domain_idx_df = df.set_index("value")
domain_idx_df.head()

Look ma, no integers! This also means that our use of .loc will have to change. Let’s try for one of my favorites:

domain_idx_df.loc["enroll-passkey.com"]

Series

Take a DataFrame and smash it apart, what would you get? Series! Each column is a Series, but don’t think of them as just glorified lists. Each Series has many of the same capabilities as a DataFrame—they even have their own Index!

We can access Series/columns a bunch of different syntaxes. We can use a dict-like square brace syntax…

df["value"]

…a list of them will work as well:

df[["value", "comment"]]

Or we can use a dot notation, if there are no spaces in the column name:

df.value

Note that these are printing with an integer column. That’s because Series also have an Index, which you normally don’t want to mess with. But it’s there!

Just take a look at everything a Series has inside of it to give you an idea of the capabilities here.

Check For Understanding

These won’t be tested, but try these challenges yourself to see if you have grasped DataFrames and the basics of data access.

  1. Use .iloc on df to access the value and comment columns of rows 30-40.
  2. Access just the comment Series using dot notation.
  3. Access the value and date Series using brace notation.

And that’ll do it for this intro to Pandas! Up next, filtering and data aggregation!

5-3: DataFrame Manipulation

Now that we have the basics of DataFrames sorted, we’re ready to use Pandas manipulate our data at scale. This is where we use grouping, aggregation

We’ll be using a log of DNS queries from Zeek, courtesy of the Security Datasets repo.

DNS logs are a great example of when we want to use Pandas to analyze aggregated data. These logs will almost always be far too large to review manually.

Our dns.log is in fact a JSON file. You might think that means we need to import the json module.

But naaaaah, Pandas has a .read_json() method. Let’s load this DataFrame up.

# Import as always
import pandas as pd

# Create our DataFrame
df = pd.read_json("dns.log")

Exploratory Data Analysis

EDA is always our first step with a proper dataset. This process gives us a general sense for the size and scope of the dataset, as well as some basic statistical information.

To start, I like to use the .shape property, which tells us the number of rows and columns as a tuple.

# (rows, cols)
df.shape

Okay, so not that big, especially for DNS logs. Let’s check the columns we have. We can do that with either df.columns or df.info(). I prefer the latter because it tells us the data type of each column.

df.info()

Ah, Zeek logs. So clean. So well-named.

These can take a while to get used to. At this point, it’s a good idea to check a few rows to see what these look like. df.head() will print the first 5 events.

df.head()

You may notice that between proto and AA columns is an ellipses. By default, Pandas will abbreviate output to make things easier to read. This is true for both rows and columns, but sometimes that’s not what we want. In this case, I really do want to see all 29 columns! I’m willing to scroll to the right!

To change this, we can use pd.set_option() to change display.max_columns to a value of our choosing. Let’s do that and re-run .head()

# Increase max cols and re-run .head()
pd.set_option("display.max_columns", 30)
df.head()

Behold! A scrollbar! Now we can see all the columns and determine where the data of interest lives.

Looks like for this dataset, id_origin_p, query, and answers have the most interesting data. That’d be the requesting IP address, the DNS query, and the responses, if any.

Why don’t we make our column names a little nicer? We can rename columns with the DataFrame’s .rename() method. It takes a dict of shape {"current_name": "new_name"} for every column you want to change. This is passed as the columns arg. And don’t forget, to make it stick in our current DataFrame, we need to use inplace=True.

I like to make sure my column names are clean right at the start of any Pandas work, and I often document the column names up front as well.

# Rename columns
df.rename(columns={"id_orig_h": "source_ip", "id_resp_h": "response_ip"}, inplace=True)
# Show results
df[["source_ip","response_ip"]].head()

Grouping and Aggregation

Ever made a Pivot Table in Excel? Not super fun, right? Turns out Pandas has similar capabilities with just a few method invocations. Very often we will want to group our data by a field. For example, what if we wanted to see how many requests each IP in our dataset generated?

Welcome to .groupby().

.groupby() takes a field or a list of fields to group by. However, the result is a little odd because it is groupby object that doesn’t show tabular data yet. That’s because we need to chain it with an aggregation function. There are many built-in functions like .count(), .mean(), and .sum(), just to name a few. Keep in mind of course that these are all quantitative functions. They have to be, because we’re talking about combining multiple rows of data into a single row. The only way a computer could sensibly do so is with mathematical functions.

Let’s group by our source IP and look at the count of how many queries each source provided.

# Group by source
df.groupby("source_ip").count()

That’s a lot of noise, since Pandas ran a count on every. single. field.

If we want to clean it up, we can specify a column at the end. But while we’re at it, I also like to chain .count() with .sort_values() to get a descending count. .sort_values() takes a by argument that tells it what field to sort by, and an optional ascending argument to toggle ascending/descending. When sorting, you want to make sure you use the field that has the count you actually want. Look closely at the table above—not every column for a given row has the same value!

# Group by source, but cleaner and descending
df.groupby("source_ip").count().sort_values(by="query", ascending=False)["query"]

Very clean! But groupby() can also take a list. Let’s try grouping by source IP and query, then count it up. We’ll use uid as our sort-by. We’ll also just display that field for cleanliness. To make it render as a proper HTML table rather than plain text, we’ll “trick” Pandas into thinking it has a list of columns by giving it a singleton list.

Oh also, since we kinda know this is going to be large, I’m going to pregame by expanding the display.max_rows setting.

# Increase max display rows
pd.set_option("display.max_rows", 100)

# Group by source AND query, count 'em, then sort descending
df.groupby(["source_ip", "query"]).count().sort_values(by="uid", ascending=False)[["uid"]]

Filtering Values

That’s a lot, right? And likely it’d be more than we need, especially when reviewing DNS queries. We can filter our DataFrame in a lot of ways. You’ve already seen the .query() method, and there’s more to it.

But there’s also the column masking method. Essentially this filters the DataFrame by applying a mask—a Series of boolean values—to a DataFrame, and only returning rows where the Series value is True.

The syntax is a little funky, so let’s go through it step by step.

First, let’s explain the mask. We can create a mask by using a boolean expression about a Series. For example, comparing .source_ip to "10.0.1.5":

df.source_ip == "10.0.1.5"

What we get back is a Series containing bools! If we put this expression inside of square braces after our DataFrame name, we’re telling Pandas to give us rows not blocked by this mask!

# Mask the df and get source_ip and query cols
df[df.source_ip == "10.0.1.5"][["source_ip", "query"]]

Working with Series Data

Sometimes, the data in a Series is a little…finicky. Many data types like str or datetime have methods for working with them using the mask method.

For example, what if we wanted to mask on strings ending with a certain value? Pandas has a .str.endswith() method that does the job, but we have to use it properly.

Pandas, why are you like this?

I know, it’s annoying. It’s because we’re never working with one value at a time. This is a fundamental shift in our thinking about data. When working with DataFrames, we have to consider what we’re doing across rows and columns.

In our count of queries by source, you may have noticed a bunch of .dmevals.local domains. These are internal. If we’re looking for external DNS queries, we can safely exclude them.

But wait. We can use .str.endswith() to find matches, but we want the inverse match! How do?

Prepending a mask expression with ~ negates it.

Let’s run that mask and save this “slice” as a new DataFrame

# Grab our external-only queries
external_queries = df[~df["query"].str.endswith(".dmevals.local")]

# Display the results
external_queries[["source_ip", "query"]]

So obviously there’s still some noise there. We don’t really need to see the microsoft.com queries.

We have the ability to combine masks with boolean-like operators, although they work differently than normal Python and and or. Instead, inside the square braces, we use the bitwise & and |.

Let’s add a negation for ending in microsoft.com to our mask.

# Grab our external-only queries, minus Microsoft
external_queries = df[~df["query"].str.endswith(".dmevals.local") & ~df["query"].str.endswith(".microsoft.com")]

# Display the first 20 results
external_queries[["source_ip", "query"]].head(20)

So there’s definitely still noise in there, but it’s a lot less than there was! And this is how we begin to parse our data to find exactly what we’re looking for.

Check For Understanding

Try some of these challenges to see if you’ve mastered grouping, aggregation, and masking!

  1. What was the average response time (.rtt) for queries to .microsoft.com domains?
  2. What was the least-common external query?
  3. What were the top 5 queries for 10.0.1.4?

In the next lesson, we’re going to dive even deeper into Exploratory Data Analysis with some more quantitative techniques!

5-4: Quantitative Analysis

Let’s get down to business. In this lesson, we’re working with an actual, factual, malware artifact. One that you’ll want to be able to manipulate easily: a packet capture. The packet capture comes from this analysis.

You won’t always have the luxury of a PCAP, but analyzing them with Jupyter gives us superpowers.

To extract data from a PCAP in Python, we use the scapy library. Let’s import that and pull in the packets with the built-in rdpcap() function.

# Import Scapy stuff
from scapy.all import *
load_layer("tls")  
# Get packets
packets = rdpcap("emo.pcapng")

Depending on the packet type, there will be different information available. The data is separated into OSI-model layers.

Let’s start by making a DataFrame of all IP packets to get general information about the TCP/IP conversations in the PCAP. To do so, we will use the .getlayer() method to retrieve the IP layer, and the .haslayer() method to look for the IP layer.

# IP packets
ip_packets = [p.getlayer(IP) for p in packets if p.haslayer(IP)]

Let’s examine the first packet to see what we’re dealing with.

ip_packets[0]

Kinda hard to read at first, but the | separates the layers of the packet. You can see that the IP layer has src, dst, sport, dport, and len data. In the case of this packet, the next layer is the DNS application data, which will contain the DNS query, among other things.

But for now, we’re just concerned with the IP layer. Now that we know the names of the properties, we can access them directly. Let’s make a list of dicts with this information to produce a DataFrame.

# Import Pandas
import pandas as pd

# Create IP data dicts
ip_data = [{"src": p.src, "dst":p.dst, "sport": p.sport, "dport": p.dport, "len": p.len} for p in ip_packets]

# Generate DataFrame
ip_df = pd.DataFrame(ip_data)
# Review the IP DataFrame
ip_df

Even without the later layers, there’s a lot we can do with this data. We can begin with some research questions.

  1. What source transferred the most bytes?
  2. What destination ports are in play?
  3. What are the external IP addresses?

Grouping and Slicing for Truth

Just as we’ve done before, we’ll by grouping our data by a field—in this case src. But instead of count(), finally, we have a reason for another aggregator. We want to add up the len field, so sum() is our choice.

Note

The numeric_only for sum() tells Pandas to only aggregate fields with numbers in it.

# Group by src and sum up len
ip_df.groupby("src").sum(numeric_only=True).sort_values(by="len", ascending=False)[["len"]]

This isn’t super informative. Instead, we can do a 2-dimensional group to see largest conversations.

# Group by src and dst
ip_df.groupby(["src", "dst"]).sum(numeric_only=True).sort_values(by="len", ascending=False)[["len"]]

That’s better. If we group by all 4 fields, we start to see conversation sizes.

We’ll need to expand our max rows to see them all.

# Expand max rows
pd.set_option("display.max_rows", 150)

# Group by src and sum up len
ip_df.groupby(["src", "dst","sport", "dport"]).sum(numeric_only=True).sort_values(by="len", ascending=False)[["len"]]

Now we can see the conversations. Looks like a lot of HTTPS traffic, which is unsurprising.

It’s a little messy to look at the conversations bidirectionally. If we want to see outbound communications, we can use the IP pattern to slice our dataframe to only those.

# Filter for internal sources only
outbound_df = ip_df[ip_df.src.str.startswith("10.")]

# Show the outbound comms grouped by src, dst, and dport. No need for sport.
outbound_df.groupby(["src", "dst","dport"]).sum(numeric_only=True).sort_values(by="len", ascending=False)[["len"]]

Much cleaner, especially with only one source!

And with that, we’ve answered our first research questions.

Assessing DNS Data

For our next trick, we want to extract the DNS data from our packets. We’ll use the same method we did before, looking for the DNS layer with .haslayer().

# Get DNS packets only
dns_packets = [p for p in packets if p.haslayer(DNS)]

DNS query data is going to live in the qd.qname property. They’re bytes, so a little .decode() is appropriate here. We can use that to build our DataFrame.

Let’s try to do it as a one-liner this time.

dns_df = pd.DataFrame([{"id": p.id, "query": p.qd.qname.decode()} for p in dns_packets])
dns_df.head()

Looking good! Now, let’s ask some questions.

  1. What were our top queries?
  2. What were the oddball queries?

Those are the common questions. Don’t sleep on the rare/oddball queries. That’s often where you’ll find evil. Never ignore the bottom of the stack.

Hopefully this is getting familiar now. We’ll group by query and aggregate with a count().

# Count of each query
dns_df.groupby("query").count().sort_values(by="id", ascending=False)

One of those sure sticks out, doesn’t it? By the way, a solid detection opportunity is for max-length DNS queries (253 chars).

But is the long DNS query actually malicious, or just weird? We can use pydig and ipwhois right from the Notebook to find out.

Note

This domain originally resolved when the course was written, but no longer has any A or CNAME records! This can happen in an investigation, and you should know how to handle it.

# Whois that weird domain's owner?
from ipwhois import IPWhois
import pydig
dig_res = pydig.query("footprintdns.com", "A")

Hmmm, something isn’t working. We need to try another record type. In fact, let’s try many, and stop when we get a hit. As a refresher:

DNS Record Types

  • A: Anchor records, straight IP
  • CNAME: Pointers to other domains
  • MX: Mail exchange
  • NS: Authoritative nameservers
  • TXT: Metadata stored in DNS

We will loop through these and then, if something is found, use that.

# Define record types
record_types = ["A", "CNAME", "MX", "NS", "TXT"]

dig_res = []

for r in record_types:
    print(f"Querying {r} records")
    dig_res = pydig.query("footprintdns.com", r)
    if len(dig_res) > 0:
        break
dig_res

Okay, nothing but nameservers. That tells us the domain is moribund, but the registration is still active. Let’s query those custom nameservers.

# Query the first nameserver
ns_res = pydig.query("ns1.footprintdns.com", "A")
print(ns_res)
# Then 
who = IPWhois(ns_res[0]).lookup_rdap(depth=1)
print(who["asn_description"])

What da—Microsoft??

Yeah, it’s some weird tracking thing they do. It looks gnarly but is in fact legitimate.

So DNS isn’t telling us much, but that in itself can be a clue! If DNS shows nothing odd, then perhaps communication was done directly via IP!

IP Analysis/Data Enrichment

Of course we have all the IP data from these communications. It’d be nice if we would add whois data like the above to each IP. And what if we could run each against VirusTotal for any information?

We can.

Let’s start with df_outbound, which handily already has our external IPs for us. We just want unique IPs, so we don’t need every row in that DataFrame. In fact, the groupby() will do nicely. We want the unique IPs, so we can export the index of the groupby(). While this will give us an Index object, we can get the raw values with the .values property.

You might notice that the result is not a list, but an array. This is an object from the numpy library. It has some more capabilities than a list, but works similarly enough for our purposes.

We’re going to build a new DataFrame column by column, which we haven’t done before. To do this it’s imperative that each list or Series that we add is the same length. We’ll base everything off our outbound_ips array.

# Get just unique destination IPs. Exclude the first entry as that's our internal IP
outbound_ips = outbound_df.groupby(["dst"]).count().index.values[1:]
# Show the IPs array
outbound_ips

Now that we have the IPs isolated, let’s build our new DataFrame. We’ll pass the constructor a slightly different object than before. Instead of a list of dicts, we’ll pass a single dict with the key as a column name, and the value as the column values.

# Begin the DataFrame with our outbound_ips
ips_df = pd.DataFrame({"ip": outbound_ips})
ips_df

Now, let’s enrich this data with RDAP lookups.

Warning

You need to handle potential errors in long-running processes. It’s also a good idea to specifically except KeyboardInterrupt during long loops. That way, the stop button always works to kill the cell.

Also, when creating Series, the size must match the DataFrame you’re adding it to. That means sometimes filling in placeholders in the event of errors. We want to account for that in our data collection.

# Initialize the list of results
whois_data: list = []

for i in outbound_ips:
    print(f"Getting {i}...", end="")
    # Use whois shell command
    try:
        whois_result = IPWhois(i).lookup_rdap(depth=1)
        whois_data.append(whois_result)
        print("Done.")
    except KeyboardInterrupt:
        break
    except Exception as e:
        print(e)
        # We have to put something in here so we have a placeholder for the Series
        whois_data.append({"asn_country_code": "None"})
        continue
# Set the `whois_country` column
ips_df["whois_country"] = [w["asn_country_code"] or None for w in whois_data]
ips_df

Look at that! Geo data! And one distinct outlier.

But we’re not done just yet. Remember back in 4-2, when we used the VirusTotal API? Let’s try searching for each of these IPs and saving the data in a new column.

We’ll start by importing what we need for VirusTotal.

# import the VT library and dependencies
import vt
import asyncio
from getpass import getpass
vt_api_key = getpass("VirusTotal API Key:")

Merging DataFrames

Sometimes working with two DataFrames is easier than one. In this case, we’re going to populate our VirusTotal data into a separate DataFrame, then merge it into the original.

We’ll start by building the new DataFrame from the ip column of ips_df.

vt_data = pd.DataFrame(ips_df["ip"])
vt_data["vt_harmless"] = 0
vt_data["vt_malicious"] = 0
vt_data["vt_suspicious"] = 0
vt_data.set_index("ip", inplace=True)

Let’s see what we made!

vt_data.head()

See how we used set_index() to make the ip column our Index? That’s going to allow us to reference rows directly by IP address. That way, we can directly set values in the columns from the VirusTotal calls. This is also why we pre-populated the columns with blank values. Now we can use those cells!

for i in ips_df.ip.values:
    # Remember VirusTotal needs async!
    async with vt.Client(vt_api_key) as client:
        r = await client.get_object_async(f"/ip_addresses/{i}")
        vt_data.loc[i, "vt_harmless"] = r.last_analysis_stats["harmless"]
        vt_data.loc[i, "vt_malicious"] = r.last_analysis_stats["malicious"]
        vt_data.loc[i, "vt_suspicious"] = r.last_analysis_stats["suspicious"]
vt_data.head()

Now, it’s trivial to combine these two with the .merge() method on the ips_df DataFrame.

Note

We do have to reset the index for vt_data. Otherwise, ip won’t be a common column!

ips_df = ips_df.merge(vt_data.reset_index())
ips_df

Finally, we’ll sort by those 3 columns in malicious, suspicious, and harmless orders to see what floats to the top.

ips_df.sort_values(by=["vt_malicious", "vt_suspicious", "vt_harmless"], ascending=False)

We now have reason to suspect that the communication with 182.162.143.56 is suspicious. We can continue our investigation with other data sources, using this as a correlation point.

6-1: Visualization

I bet by the end of last lesson, you were tired of looking at giant tables. Sometimes tabular representation is the right call to tell the story of your data, but other times, a picture is worth more than ten thousand table rows.

And when you have that many, it’s definitely best to avoid looking at them all row-by-row. To prove that point, we’re going to work a very large file of many log entries: in this case, Sysmon logs from a Conti Ransomware event. It’s so big, in fact, we can’t include the file in the repository. So let’s download it now.

# Download logs
! curl -L -o windows-sysmon.log https://media.githubusercontent.com/media/splunk/attack_data/master/datasets/malware/conti/conti-cobalt/windows-sysmon.log

Parsing Sysmon Logs

There are myriad PowerShell tools that will parse Sysmon logs and allow incident responders and threat hunters to quickly, efficiently review event data. But we’re not in PowerShell, are we? Unfortunately, I was unable to find anything in the Python universe to do the same.

So I wrote one.

The code in sysmon.py is part of my own library of tools. Feel free to read through it, modify it, and expand it! What this module does is deconstruct WinEvent XML with BeautifulSoup to produce SysmonEvent objects, each with useful properties we can access and compare.

There’s also a handy load_xml() function that will load the raw XML directly into a list of SysmonEvent objects. There’s another load_evtx() function for files in the binary EVTX format, but that’s not what we have here.

This file will be our starting point for this lesson. Let’s get those events!

# Import Sysmon module and Pandas, because of course
import sys
sys.path.append("../lib/")
import sysmon
import pandas as pd
# Load Sysmon Events from the `windows-sysmon.log` file.
evts = sysmon.load_xml("windows-sysmon.log")

If all went well, we should have a nice list of events. Before even loading this into a DataFrame, let’s look at how many events we have.

# len of evts
len(evts)

So that’s…a lot. But what does one of these things look like, anyway? Let’s examine one of them

# Look at the first Event object
evts[0]

That could be clearer. What does the object contain? One easy way to find out is with the built-in vars() function, which will extract the attributes of an object to a dict.

# dict-ify our evt
vars(evts[0])

There now, that’s something. While different Event IDs will contain some different properties, each one contains:

  • user_id
  • event_id
  • time_created
  • pid

We can use vars() to quickly make a big ol messy DataFrame.

You might ask why we didn’t do that in the first place. Sysmon Events are complicated enough that having the structure of these objects becomes more valuable in the long run. For now, let’s try to make that DataFrame by giving Pandas the vars() result of our objects.

We’re also going to drop the soup field because we don’t need all that data, and it’s gigantic.

df = pd.DataFrame([vars(e) for e in evts])
# Drop the sop column because we don't need it and it's huge
df.drop("soup", axis=1, inplace=True)
df

There now, a proper DataFrame. But that’s a table and this lesson is about visualizations! What kind of visualizations can we do? What about, say, a chart of Event IDs?

Easy enough, right? Kinda!

The good news is DataFrames have a built-in plot property that exposes all kinds of graphing features. They take some getting used to though, and they also require importing the matplotlib library. Let’s take take of that before we do anything else.

# Import matplotlib
import matplotlib

matplotlib is basically an entire course, and I will not be going into every aspect of it. We’ll take what we need to make basic visualizations with it.

To start, let’s say we want a simple bar chart, showing the count of every Event ID. It seems simple, but remember that right now our DataFrame has a row for every Event. It is not yet aggregated in any meaningful way. So up first, we know we’ll have to groupby() and, since we want how many events per, we’ll use .count() as our aggregator. We’ll even sort them for ease of use.

For ease, we’ll save this slice as a new variable.

# Save the counts as a new slice. This is common when working with the data
# in multiple ways
event_counts = df.groupby("event_id").count().sort_values(by="pid", ascending=False)
event_counts

Now to move from a table to a graph. The built-in plot.bar will provide us a bar graph of this data!

Using the plot methods takes some practice. Most of the arguments it requires are named, not positional, so you need to be familiar with the spec. I always have the documentation open when making these.

At base, the plot needs x and y values. The x axis is how we’re grouping our data (event_id) in this case, and the y axis is our value to measure. Since we count()ed everything, we want to use a column that is present on all rows. pid will do nicely.

Let’s try just that much.

event_counts.plot.bar(x="event_id", y="pid")

What?? What went wrong? KeyError: 'event_id'? How can that be?

Welcome to Pandas. Because we grouped by event_id, it is no longer an accessible column; it’s the Index. The very thing we needed to do to shape our data removed the column we need for the chart.

But worry not, there’s a solution: we can reset_index() to make the Index a RangeIndex once more and get our event_ids column back. Our data shape is the same, but now we have the column we need. One more time!

# Once more, but with a reset index
event_counts.reset_index().plot.bar(x="event_id", y="pid")

We made a thing! It could be prettier, but it’s a chart nonetheless. matplotlib defaults aren’t too shabby. In the next lesson we’ll go over some more advanced graphing libraries, but there’s more we can do to make this helpful. For one thing, every chart needs a title. Also, can we do something about those axes labels? Basic middle school science stuff.

The easiest way to make these modifications is to save the chart to a variable and work with it there. You can see above the chart that this is an AxesSubplot object. Commonly we’ll name Axes objects ax. We can actually set the title with an argument to plot.bar(), but the label and the legend need methods on the axes.

# Safe the axes object to manipulate it
ax = event_counts.reset_index().sort_values(by="event_id").plot.bar(x="event_id", y="pid", title="Count of Event IDs")
ax.set_xlabel("Event ID")
ax.set_ylabel("Event Count")
ax.legend(["Events"])

Much better! Now that we have the basics down, let’s split our DataFrame a bit more to examine our data.

Process Executions

Event ID 1 in Sysmon is a goldmine. Let’s begin investigating process executions in our logs by isolating that Event ID

# Get all process execs
process_execs = df[df.event_id == 1]
# Get the shape of the new df
print(process_execs.shape)
print(process_execs.columns)
process_execs.head()

Awesome, 229 events to work with. How might we visualize this data?

Perhaps as a starter we can look at the image field for rare occurrences. This can be done in tabular as well as graphical format. Let’s get both going.

# Group by image
procs_by_image = process_execs.groupby("image").count().sort_values(by="pid", ascending=False)
procs_by_image

I love this dataset; it’s so clean. Real life doesn’t always work like that. But even here, we can see some sus-looking images. beacon.exe indeed.

Does a chart of this information help us? I don’t think it’s super valuable as-is, but let’s make one to see it anyhow.

ax = procs_by_image.reset_index().plot.bar(x="image", y="pid", title="Executions by Image")
ax.set_xlabel("Image")
ax.set_ylabel("Execution Count")
ax.legend(["Executions"])

Which way did you tilt your head to look at that—left or right? The answer says a lot about you!

Didn’t tilt your head at all? Okay, you psychopath.

But seriously, this isn’t a great way to display this data. In the case of process executions, a table is better. Why did we do this at all then? To drive home this point:

Use the right form to tell your data’s story. Tables aren’t always correct and neither are charts.

Network Connections

This one is going to lend itself much better to visualization. Event ID 3 shows network connections, and has source, destination, and port information in there. Let’s grab ’em.

We’ll also just print out the columns we care about.

# A new slice for EventID 3
network_conns = df[df.event_id == 3]
print(network_conns.shape)
print(network_conns.columns)
network_conns[["time_created","pid","image","user","src_ip","dest_ip","src_port","dest_port"]].head()

Ports

Let’s see what ports were most used by grouping and graphing. We’re going to do this all in one move. The line of code gets long, but we can split it with \ characters.

Oh one other trick. We are going to only take the top 10 entries with a creative sort. Why? Because without that, the number of ports would make this utterly unreadable.

# Group and graph
ax = network_conns \
  .groupby("dest_port") \
  .count() \
  .reset_index() \
  .sort_values(by="pid", ascending=False) \
  .head(10) \
  .plot.bar(x="dest_port", y="pid", title="Network Connections by Destination Port")

ax.set_xlabel("Dest Port")
ax.set_ylabel("Conn. Count")
ax.legend(["Connections"])

Now that we have an overview, we can dive in on the points of interest in the dataset.

Zooming In on Processes of Interest

Now we see some interesting processes right? beacon.exe in particular. We can do some really interesting graphing/tabling with that information.

Let’s isolate that in a new slice and see what kind of events it emits. We’ll do so with the image and parent_image properties, looking for what the process did, but also anything its immediate child processes did. Then we’ll graph the events.

# Slice for beacon or beacon's kids
beacon_exe = df[df.image.str.endswith("beacon.exe") | df.parent_image.str.endswith("beacon.exe")]
# Graph Events
ax = beacon_exe.groupby("event_id").count().reset_index().plot.bar(x="event_id", y="pid", title="Beacon.exe Event IDs")
ax.set_xlabel("Event ID")
ax.set_ylabel("Event Count")
ax.legend(["Events"])

So a few process execs, a lot of network activity, some file writes? and a massive amount of registry activity. That’s not exactly uncommon, but worth noting.

It’s important to see that there are in fact network connections. To what?

# Beacon Network Connections
ax = beacon_exe[beacon_exe.event_id == 3] \
  .groupby("dest_ip") \
  .count() \
  .reset_index() \
  .plot(kind="bar", x="dest_ip", y="pid", title="Beacon.exe Destination IPs")
ax.set_xlabel("Dest IP")
ax.set_ylabel("Conn. Count")
ax.legend(["Beacon.exe"])

ax = beacon_exe[beacon_exe.event_id == 3] \
  .groupby("dest_port") \
  .count() \
  .reset_index() \
  .plot(kind="bar", x="dest_port", y="pid", title="Beacon.exe Destination ports")
ax.set_xlabel("Dest Port")
ax.set_ylabel("Conn. Count")
ax.legend(["Beacon.exe"])

Mostly on 443. Unsurprising, but that does mean we may have one other bit of data to grab: DNS, aka Event ID 22.

I think this may be nicer as a table.

# Check out event ID 22
beacon_exe[beacon_exe.event_id == 22][["time_created","query_name","query_results"]]

Well hello there. Here we have a C2 calling home to multiple subdomains, but all resolving to the same IP.

There’s plenty of other investigation we can undertake with this sample, but much of it may require a bit more complex visualization. That’s why we’re going to stop here for this lesson.

Take some time to practice with df.plot. Try other chart types! Try stacked plots!

In the next lesson, we’ll explore more complex visualization libraries. See you there!

6-2: Advanced Visualization

Our last lesson introduced the concepts around visualization: x and y axes, plotting, and slicing DataFrames specifically for use with graphing. You probably wished that the charts we made were better-looking, or bigger, or dynamic, or all of the above.

Me too. Even though Matplotlib is great for a lot of things, here’s the truth: I never use it for my charts in real life. It’s just too inflexible.

Instead, I use Plotly.

Plotly is my first stop. For most of the charts I need to make, Plotly is flexible and visually appealing. Altair does one thing much, much better though which makes it indispensible: time series.

We’ll get to how those work in a moment. For now, let’s import this library and see how it differs from Matplotlib.

Even importing Plotly is a bit different. We pull a submodules: plotly.express. Sometimes we also need plotly.graph_objects, but not for what we’ll be doing here

# Import our libraries
import sys
sys.path.append("../lib")
from sysmon import load_xml
import pandas as pd
import plotly.express as px

Now, let’s get our next sample. It’s more Sysmon data, but from a different incident.

# Download the Sysmon data.
import os
if not os.path.exists("sysmon.log"):
    ! curl -L -o sysmon.log https://media.githubusercontent.com/media/splunk/attack_data/master/datasets/malware/hermetic_wiper/sysmon.log

And once more, let’s load our DataFrame.

# Create our DataFrame once more
evts = [vars(e) for e in load_xml("sysmon.log")]
df = pd.DataFrame(evts)

Okay, let’s get comfortable with a basic bar chart: a count of Event IDs. In Plotly, we pass the DataFrame, the x and y fields, and a title. We can also pass a label argument for the axes. That’s pretty much it. Let’s do it!

Don’t forget we need the values to be correct, so we’ll do groupby() to get our counts.

# Create our bar chart
events_by_id = df.groupby("event_id").count().reset_index()
px.bar(events_by_id, x="event_id", y="pid", title="Count of Event IDs", labels={"event_id":"Event ID", "pid":"Event Count"})

So it’s not perfect, but look at what we have. In addition to a much larger viewspace, the chart is interactive! We can hover over the bars for more information, zoom in/out, and even download the chart as a .png. This is why I use Plotly.

But this is hardly as nice as we want the chart. For one thing, we’d want to fix the X axis. Plotly interpreted the Event IDs as numbers, and therefore interpreted it as a range. We can fix that by forcing the axis to be categorical, rather than numerical data. To do this, we will create a stringified Event ID column and use that for our X-axis and our color!

# Add stringified event_id
events_by_id["event_id_str"] = events_by_id.event_id.astype("str")
# Same chart, but with a categorical, stringified x-axis, and color!
px.bar(events_by_id, x="event_id_str", y="pid", title="Count of Event IDs", labels={"event_id_str":"Event ID", "pid":"Event Count"}, color="event_id_str")

Not too shabby, right? Resize your browser window.

Aww yiss, responsive charts!

Now let’s see what else we can do.

Multi-dimensional analysis

Right now we’re plotting x and y, which is handy of course. But what if we could add a 3rd dimension? The color gives us this capability.

For example, what if we looked at network connections by destination IP. And then, what if we could split them by image to identify what processes were making the most noise?

Instead of stringifying the dest_port, we will modify the figure not unlike our modification of the axes in Matplotlib. We can use update_xaxes to set the axes type as categorical, making Plotly track the dest_port correctly.

Now, we need counts of connections by both dest_port and pid, but how do we do that?

If we groupby() both and then reset the Index, we’ll get both values in individual rows, along with the counts. This will provide us both values at once to graph. We can count user_id as our y-axis.

# Slice our df to get network connections, then group by out 2 values
network_conns = df[df.event_id == 3].groupby(["image", "dest_port"]).count().reset_index()
fig = px.bar(network_conns, x="dest_port", y="user_id", color="image", labels={"dest_port": "Destination Port", "user_id": "Connection count"}, title="Connections by Dest Port/Image")
fig.update_xaxes(type="category")
fig.show()

This technique is a powerful view into event data. Being able to cut across 2 dimensions at once reveals outliers and correlations.

Time is a tricky thing. Python represents time data in datetime object. Right now, the DataFrame’s time_generated field is just a str. Luckily, it’s of such a format that we can easily convert it to a datetime object with Pandas’s built-in to_datetime() method. We’ll add that as the _time field, just like a Splunk dataset.

# Create a proper datetime field
df["_time"] = pd.to_datetime(df.time_created)

# Drop the soup field to avoid weird issues with Plotly
df.drop("soup", axis=1, inplace=True)
# Examine our new column
df._time

Having the _time field by itself isn’t all that helpful. We’d need to get counts of events grouped by time in some way. Unfortunately, Plotly can’t do that for us. We need to use the datetime methods to group our data usefully.

pd.Grouper() allows us to use the datetime as a grouping mechanism. The function takes a key for the datetime field to use, and a freq argument that defines the grouping frequency. The freq argument takes a specially-formatted string that indicates the grouping interval. Let’s give it a try in tabular form to see the results.

# Group by time, every 10s
df.groupby(pd.Grouper(key="_time", freq="10s")).count().head()

So! Our Index (which we will need to reset for graphing) is a timestamp, but notice: each row is 10 seconds later. Our events have been grouped and counted in those time, uh, buckets. Time buckets. That’s what I’m going with.

Plotly can easily use these time buckets to plot events. Let’s do a raw events over time with this method.

# Raw events over time
events_10s = df.groupby(pd.Grouper(key="_time", freq="10s")).count().reset_index()
px.line(events_10s, x="_time", y="pid", title="Events Over Time, 10s interval", labels={"_time":"Time", "pid": "Event Count"})

Change the freq value to see how the bucket size changes the story we’re telling. At the 10s size, it seems like something interesting was happening around 12:25:30. A smaller bucket (finer resolution) gives us even more precision.

Now about those colors. What if we wanted to graph the Event IDs separately?

We can use the same trick with used with dest_port and image—a double .groupby()—to split these by Event ID.

# Events by Event ID
events_2s = df.groupby([pd.Grouper(key="_time", freq="2s"), "event_id"]).count().reset_index()
px.line(
    events_2s, 
    x="_time", 
    y="pid", 
    title="Events Over Time, 2s interval", 
    labels={"_time":"Time", "pid": "Event Count", "event_id": "Event ID"}, 
    color="event_id"
)

This really tells an interesting story about what was happening in the 12:25:30 time space. Even though this is just raw volumetric analysis, we see spikes across almost all Event IDs.

Oh, did you click on items in the legend yet? You can show/hide different series.

I like line charts for time series, but a stacked bar can also give us a nice sense of proportionality. Let’s redo the same chart, but use a bar type instead. For whatever reason, to make the colors work, we’ll need to stringify the event_id.

# The same, but bar style
events_2s["event_id_str"] = events_2s.event_id.astype("str")
px.bar(
    events_2s, 
    x="_time", 
    y="pid", 
    title="Events Over Time, 2s interval", 
    labels={"_time":"Time", "pid": "Event Count", "event_id_str": "Event ID"}, 
    color="event_id_str"
)

I think it’s cool how, especially when you zoom in on the 12:25:30 section, you can really see the proportionality of different events.

Non-Standard Datetimes

Every so ofen you’ll come across a log type that has weirdo timestamps that Pandas can’t parse. Like, I dunno, SYSLOG?!

# Totally reasonable Syslog format that Pandas can't handle
try:
    pd.to_datetime("Nov 13 09:03:52")
except:
    print("Pandas can't do that for some reason ¯\_(ツ)_/¯")

Luckily, pd.to_datetime() has a format optional argument that takes the formats supported by strftime.

So we can in fact teach Pandas how to read Syslog datetimes.

# One more time, with better information
pd.to_datetime("Nov 13 09:03:52", format="%b %d %H:%M:%S")

Closer, but we don’t have year information so now it’s…1900?

Pandas has a cool built-in offset library that allows us to fix things like this. The DateOffset object can be added to the datetimes to fix our years.

# Now we're in the present!
pd.to_datetime("Nov 13 09:03:52", format="%b %d %H:%M:%S") + pd.offsets.DateOffset(years=122)

6-3: Reporting

Not everyone wants to read all your code, nor should they. But the information your Notebooks produce will be valuable to a wide range of audiences. So how to best provide the information?

Let’s start by grabbing some logs. In this case, HTTP logs from Splunk’s Boss of the SOC, v1. We also need to clean up the data for use in Pandas, which this will handle.

import os

if not os.path.exists("./iis.json"):
    ! wget https://s3.amazonaws.com/botsdataset/botsv1/json-by-sourcetype/botsv1.iis.json.gz -O iis.json.gz
    ! gunzip iis.json.gz
    ! sed -i '1s/^/[/' iis.json
    ! sed -i 's/}$/},/' iis.json
    ! sed -i '$s/,$//' iis.json
    ! echo "]" >> iis.json

Now, let’s make some fun charts!

import pandas as pd
import json
import plotly.express as px

# Mute an annoying PerformanceWarning
from warnings import simplefilter
simplefilter(action="ignore", category=pd.errors.PerformanceWarning)

with open("iis.json") as f:
    json_evts = [j["result"] for j in json.load(f) if "result" in j.keys()]

df = pd.DataFrame(json_evts)
df.shape
df.head()
http_method = df.dropna(subset=["cs_method"]).groupby("cs_method").count().reset_index()
px.bar(http_method, x="cs_method", y="c_ip", labels={"cs_method": "HTTP Method", "c_ip": "Request count"}, title="Requests by HTTP Method", color="cs_method")

# Setup datetime
df["_time"] = pd.to_datetime(df._time.str.replace("MDT", "-0700"))
methods_by_time = df.groupby([pd.Grouper(key="_time", freq="1D"), "cs_method"]).count().reset_index()
px.bar(methods_by_time, x="_time", y="c_ip", labels={"_time": "Time", "c_ip": "Request Count", "cs_method": "HTTP Method"}, color="cs_method", title="HTTP Requests by Day")

# Create time boundary
aug_10 = df[df._time.dt.day == 10]
aug_10_10m_methods = aug_10.groupby([pd.Grouper(key="_time", freq="10min"), "cs_method"]).count().reset_index()
px.bar(aug_10_10m_methods, x="_time", y="c_ip", labels={"_time": "Time", "c_ip": "Request Count", "cs_method": "HTTP Method"}, color="cs_method", title="HTTP Requests by Day")

aug_10_s_ip = aug_10.groupby("c_ip").count().reset_index()
px.bar(aug_10_s_ip, x="c_ip", y="s_ip", labels={"c_ip": "Client IP", "s_ip": "Request Count"}, color="c_ip", title="8/10/16: Requests by Client IP")

Okay so we have all of this data, and maybe more. How do we report it?

DataFrames have built-in .to_csv() and .to_excel(), but that’s just tabular data. What about providing data to leadership?

Jupyter has the ability to export data to PDF as well. However, that’s not my preference because while it works, it’s not the most visually appealing.

Instead, we can export an HTML page that contains the charts, but omits the code. We’re assuming here that the audience is interested in the output. If that’s not so, leave the code in—although at that point they might just want the Notebook.

To generate a code-free HTML from a Notebook, we leverage the jupyter command line, like so.

! jupyter nbconvert --to html --no-input 6-3_reporting.ipynb

Refresh the file browser on the side to view the HTML.

Jupyter has a few other export options as well, including a slideshow! Click on the gear icon on the right side to find a “Slide Type” setting. This allows you to choose whether a cell is a Slide by itself, a Fragment (appears later), a Sub-Slide (drill-down), Skip (omit), or Notes. The result is a bit wonky, but handy for presenting visual data.

Reporting Alternatives

There are two other tools I want you to be aware of for reporting. The first is Mercury, a comprehensive platform to turn Jupyter Notebooks into interactive applications for non-technical audiences. The second is Marimo, an alternative to Jupyter entirely that makes the entire Notebook out of a Python file. The editing experience is similar to Jupyter, but with some important differences and creature comforts. There is also an “App Mode” that allows a clean display of data in a web interface. Marimo is also, for better or worse, deeply inegrated with AI tools.

Regardless of tool, your work in Notebooks should be presentable to audiences that don’t need to care about the code. The code is for repoducibility and ediitng by the creators, but not by leadership whose decisions your analysis informs.

The End

Aaaand…that’s it!

This is the Last Notebook. You made it! Congratulations!

If you’ve practiced the skills and techniques in this course, you’ll be ready for the Exhibition of Mastery. There’s one more unit on Next Steps, then: the Exhibition!

Congratulations, and thank you for your time, attention, and support of the Taggart Institute. I hope you’ve found the course worthwhile.

- Michael

Exhibition of Mastery