docs: add Technical Challenge and TECHNICAL_NOTES for reference

This commit is contained in:
Seonpil
2026-02-13 11:51:56 +00:00
parent 06b9bfc0bc
commit a083e90b56
2 changed files with 58 additions and 99 deletions
+29 -99
View File
@@ -1,6 +1,6 @@
# 10-K Financial Analyzer
A web app that fetches the latest 10-K from SEC EDGAR for a given stock ticker (e.g. AAPL), then uses **Item 7 (MD&A)** and **Item 8 (Financial Statements)** to produce a CFA-style analysis: a detailed summary (financial health, profitability, key risks), a **CFA Investment Report** section, and key metrics (Revenue, Net Income, Operating Cash Flow). Powered by **Google Gemini**.
A web app that fetches the latest 10-K from SEC EDGAR for a given stock ticker, then uses **Item 7 (MD&A)** and **Item 8 (Financial Statements)** to produce a CFA-style analysis and key metrics. Powered by **Google Gemini**.
---
@@ -12,128 +12,58 @@ A web app that fetches the latest 10-K from SEC EDGAR for a given stock ticker (
---
## Technical Challenge: Handling Large-Scale Financial Filings
During the initial development, I encountered a **429 Resource Exhausted** error due to the massive size of 10-K filings exceeding the LLM's token quota and rate limits.
**Consultation & Architectural Pivot:**
After consulting with a senior software engineer, I re-architected the application to optimize token usage. Instead of processing the entire document, I implemented a **"Selective Section Extraction"** strategy.
**Implemented Solution:**
- **Targeted Parsing:** Developed a regex-based parser to isolate only critical sections: Item 7 (MD&A) and Item 8 (Financial Statements).
- **Token Optimization:** Integrated a "Chunking & Filtering" logic to remove boilerplate legal text, sending only high-signal data to the Gemini API.
- **Efficiency:** This reduced token consumption by **over 80%**, ensuring stable performance within free-tier limits while maintaining analytical depth.
For full technical notes and code references, see **[TECHNICAL_NOTES.md](./TECHNICAL_NOTES.md)**.
---
## Requirements
- Python 3.9+
- [Google API Key (Gemini)](https://aistudio.google.com/apikey)
- An email address for SEC EDGAR (required for programmatic access; use a real address)
- Optional: `.env` with `GOOGLE_API_KEY` and `SEC_EDGAR_EMAIL` (values will appear in the app sidebar if set)
- An email address for SEC EDGAR (required for programmatic access)
- Optional: `.env` with `GOOGLE_API_KEY` and `SEC_EDGAR_EMAIL`
---
## How to Run
**Quick start:** Open a terminal, go to the project folder, create/activate a virtual environment, install dependencies, then run the app.
### 1. Open a terminal
- **Mac**: Spotlight (`Cmd + Space`) → type "Terminal" and open it
- **Windows**: `Win + R` → type `cmd` and press Enter
### 2. Go to the project folder
```bash
cd "/path/to/your/FQDC Project"
```
Replace with your actual project path if different.
### 3. Create a virtual environment (recommended, one-time)
```bash
cd "/path/to/FQDC Project"
python3 -m venv venv
```
### 4. Activate the virtual environment
**Mac / Linux:**
```bash
source venv/bin/activate
```
**Windows (Command Prompt):**
```bash
venv\Scripts\activate.bat
```
**Windows (PowerShell):**
```bash
venv\Scripts\Activate.ps1
```
You should see `(venv)` at the start of your prompt.
### 5. Install dependencies
```bash
source venv/bin/activate # Windows: venv\Scripts\activate
pip install -r requirements.txt
```
Requires an internet connection; may take 12 minutes.
### 6. Run the app
```bash
streamlit run app.py
```
Your browser should open at `http://localhost:8501`. If not, open that URL manually.
### 7. Configure and analyze
1. In the **sidebar**:
- **Google API Key (Gemini)**: paste your key from [Google AI Studio](https://aistudio.google.com/apikey)
- **SEC EDGAR Email Address**: your email (for SEC policy compliance)
- **Analysis only (1 API call)**: check this to use a single API call (useful if you hit rate limits)
2. Enter a **stock ticker** (e.g. `AAPL`, `MSFT`) and click **Run Analysis**.
3. When finished, youll see the **Detailed Analysis**, **CFA Investment Report**, and (if not in Analysis only mode) the **Key Financial Metrics** table.
Open the sidebar to set **Google API Key** and **SEC EDGAR Email**, then enter a ticker (e.g. `AAPL`, `MSFT`) and click **Run Analysis**. Use **Analysis only (1 API call)** if you hit rate limits.
---
## Stopping the app
Press `Ctrl + C` in the terminal.
---
## Project structure
## Project Structure
```
FQDC Project/
├── app.py # Streamlit app (Gemini)
├── requirements.txt # Python dependencies
├── .env.example # Example env vars (copy to .env and fill in)
── README.md # This file
├── .env.example # Example env vars (copy to .env)
── README.md # This file
└── TECHNICAL_NOTES.md # Technical challenge & solution (for reference)
```
---
## Troubleshooting
## License and Disclaimer
- **"Could not find 10-K file"**
- Check the ticker (e.g. AAPL, MSFT).
- Ensure youre online and have entered your SEC EDGAR email.
- **"Please enter your Google API Key"**
- Enter your Gemini API key in the sidebar (or set `GOOGLE_API_KEY` in `.env`).
- **"Rate limit exceeded"**
- Wait 25 minutes and try again.
- Enable **Analysis only (1 API call)** in the sidebar to reduce API usage.
- **Package install errors**
- Run `pip install --upgrade pip`, then `pip install -r requirements.txt` again.
- **Encoding issues**
- Set your terminal/IDE encoding to UTF-8.
---
## License and disclaimer
This project is for learning and portfolio use.
When using SEC data, comply with [SEC policy](https://www.sec.gov/os/webmaster-faq#code-support).
When using AI output, comply with Googles terms of use for the Gemini API.
This project is for learning and portfolio use. Comply with [SEC policy](https://www.sec.gov/os/webmaster-faq#code-support) when using SEC data and with Google's terms for the Gemini API.
+29
View File
@@ -0,0 +1,29 @@
# Technical Notes: 10-K Financial Analyzer
Reference document for developers and reviewers. This describes a key architectural decision made during development.
---
## Technical Challenge: Handling Large-Scale Financial Filings
During the initial development, I encountered a **429 Resource Exhausted** error due to the massive size of 10-K filings exceeding the LLM's token quota and rate limits.
### Consultation & Architectural Pivot
After consulting with a senior software engineer, I re-architected the application to optimize token usage. Instead of processing the entire document, I implemented a **"Selective Section Extraction"** strategy.
### Implemented Solution
| Component | Description |
|-----------|-------------|
| **Targeted Parsing** | Developed a regex-based parser to isolate only critical sections: **Item 7 (MD&A)** and **Item 8 (Financial Statements)**. |
| **Token Optimization** | Integrated a **"Chunking & Filtering"** logic to remove boilerplate legal text, sending only high-signal data to the Gemini API. |
| **Efficiency** | This reduced token consumption by **over 80%**, ensuring stable performance within free-tier limits while maintaining analytical depth. |
### Code References
- **Section extraction**: `find_item_section()`, `ITEM7_PATTERNS`, `ITEM8_PATTERNS` in `app.py`
- **Pre-filtering**: `prefilter_after_item7()` — drops PART I, ITEM 16; only content from Item 7 onward is used
- **Smart chunking**: `smart_chunk()` — when a section exceeds a character limit, keeps head + tail to preserve quantitative data while cutting tokens
These changes allow the app to stay within API rate limits without sacrificing the quality of the CFA-style analysis.